sankalp had never worked on GPU kernels professionally. Fourteen days and over 1,500 submissions later, they finished 12th of 183 in GPU MODE’s qr_v2 contest. That is 232x faster than the torch.geqrf baseline, on a $220 subscription stack.
Then the thread turned on the podium. One commenter says 8 of the 10 top solutions broke on any input outside the contest shapes.
In today’s indie hacker news:
- 🏁 Codex hill-climbed a GPU kernel to 232x
- 📓 AI’s math edge may be a bigger notebook
- 📏 A tape measure caught what BMI missed
- 🧪 25 proteins promised 26%, the trial delivered nothing
- 💬 HN spent 184 comments fighting one word
TOP STORIES
🏁 OVERFIT AND FURIOUS

A three file harness beat the reference implementation, and the podium code mostly did not generalize.
The story: sankalp’s writeup is a log of an agent loop, not a kernel tutorial. The problem was shaped for one: pass or fail correctness, shape-wise timings, unlimited submissions. Steering came from three markdown files. One held the problem statement, one held submission instructions, and one logged every attempt with its result.
The stack was consumer grade. A $200 ChatGPT Pro plan, a $20 Claude Pro plan, and Modal for profiling. Modal hands out $30 of free credits a month. Codex won the slot because it does not quit. Claude kept announcing it had exhausted all optimizations, sankalp says.
The details:
- The beam fix. Stuck between 3,000 and 1,800 microseconds, sankalp had the agent hold 3 to 5 live candidates, not one. A structurally new idea usually loses its first round and gets killed early.
- The template. The loop copies karpathy/autoresearch, where the human edits one spec file and the agent edits one training file. Its README caps each run at five minutes, roughly 100 experiments while you sleep.
- The receipts folder. By the end the tree held 560 submission variants, 119 Modal probe scripts, and 68 experiment writeups. The logging existed so the next session would not re-run a dead idea.
- The public score lied. Michael Lutz placed 5th and logged one route that passed every public run. It scored 2.59 ms on the public set and 9.81 ms on the secret one.
- The next contest. sankalp says most top solutions in the follow-up Cholesky problem hit about 4 of 8 on a real validation run. Not very numerically stable, in their words.
“A specific objective gave the agent room to iterate. A scalar score gave it room to cheat the intent.” (Michael Lutz, 5th place)
Why builders care: If you already own a verifier, a test suite or a benchmark, aim one loop at it. Write the objective as a contract with zero regression elsewhere, and build the held-out check before you trust the number.
Keep that gap in mind. It shows up again below, in a drug trial.
📓 BIGGER NOTEBOOK, SAME BRAIN

Davide Piffer says AI out-remembers mathematicians, and the n-back papers disagree.
The story: Piffer’s essay makes one bolded claim. AI has access to a vastly larger working memory than the human brain. His case is that the context window is an external symbolic workspace. The text is not a report of the thinking, he argues, it is part of the mechanism.
The human number he is arguing against is small. Nelson Cowan’s review puts the central store at 3 to 5 meaningful items in young adults. Strip out sensory memory and it lands near 4 items. Against that, Claude Sonnet 4 ships a 1 million token window.
The details:
- His own caveat undercuts the headline. Advertised context length is not the same as perfectly usable memory, Piffer writes. Models overlook relevant information and lose track of details.
- The measured limit points the other way. An AAAI-24 paper ran verbal and spatial n-back tasks on ChatGPT. Its finding: a working memory capacity limit strikingly similar to that of humans.
- And there is a mechanism. A 2024 follow-up trained decoder-only transformers on the same tasks. Attention entropy rises as N goes up, which is where the ceiling comes from.
- Filling the window is a line item. Input pricing on that window steps from $3 to $6 per MTok above 200K tokens. Output goes from $15 to $22.50.
- The reach check. The essay carries 44 likes and 13 restacks on a newsletter listing over 2,000 subscribers. It landed on Hacker News with 441 points.
“What looks like deeper thought may sometimes be broader search conducted inside a much larger notebook.” (Davide Piffer)
Why builders care: If the edge is bookkeeping, the cheap win is a better ledger, not a bigger model. Restate the live constraints every step, force intermediate state into text, and wire in a checker.
📏 BMI HAD ONE JOB

A tape measure reclassified heart risk that BMI had already cleared.
The story: The American College of Cardiology announced a JACC paper from the Cross Cohort Collaboration. It tracked over 260,000 people for an average of 20 years across nine cardiovascular outcomes. Inside the group BMI called normal weight, 5% had a high waist circumference and 18% had a high waist-to-hip ratio.
The misclassification runs both directions. Among people BMI called obese, 45% carried a low waist-to-hip ratio. Obesity by BMI with a low waist showed no significantly different risk from normal weight.
The details:
- The risk gap. Normal-weight or overweight people with a flagged waist carried 15% to 50% more risk on most of the nine outcomes.
- Adjusting for BMI makes the waist signal stronger. In EPIC, the top waist quintile carried a 1.33 hazard for death before BMI adjustment and 2.05 after.
- Your own tape is accurate enough. Self-measured and technician-measured waist correlate at 0.8 to 0.9. Self-measures run about 1 cm to 3 cm low.
- The dose is known and it is small. In a 300-person trial, every exercise arm cut waists about 5 cm versus control. Doubling the exercise did nothing extra.
- The honest caveat. Adding waist circumference to the Framingham Risk Score or Pooled Cohort Equations has not improved discrimination. This is a personal baseline, not a new screening score.
Why builders care: The desk-bound founder with a clean smart-scale readout is exactly the profile this study says gets missed. John Molvin’s rule fits on one line and costs nothing: waist under half your height.
🧪 THE SCORE MOVED, THE PATIENTS DIDN’T

A protein panel predicted a dementia benefit that the randomized trial never produced.
The story: Alzforum’s CTAD coverage puts the two results side by side. Novo Nordisk scored plasma samples from the SELECT trial with a machine learning panel called the dSST. Its input is 25 proteins covering neuroinflammation, immune dysfunction, synaptic integrity and metabolism. Two years of semaglutide scored a 26 percent lower predicted five year dementia risk, and 9 percent at 20 years. The poster itself sits behind a healthcare professional login.
Those are model outputs, not diagnoses. The randomized readout ran in 3,808 patients across 40 countries. On the primary CDR-SB endpoint at week 104, the curves sat on top of each other. Every secondary cognitive and functional outcome matched.
The details:
- The drug did things, just not that thing. Plasma hs-CRP dropped about 30 percent. Participants lost 5.8 percent of body weight, against a 0.6 percent gain on placebo.
- The blood markers went the wrong way. CSF markers fell by 10 percent or less. In plasma, GFAP rose 4 percent and NfL rose 5 percent, which the presenter could not explain.
- The drug barely reached the brain. Median semaglutide sat at 0.14 nmol/L in CSF against 34.8 nmol/L in plasma. That is roughly 0.4 percent.
- The evidence that greenlit it was observational. A 1,710,995 patient EHR study reported a 0.54 hazard ratio against insulin. Novo Nordisk skipped a dose finding Phase 2 entirely.
- Not everyone accepts the null. Christian Hoelscher argues the trial analyzed one timepoint and ignored later ones. His reading comes from presented slides, not a published analysis.
Why builders care: This is the cleanest available case study in proxy metric drift. If your roadmap runs on a score your own model emits, run the controlled test before you scale the bet.
💬 THE WORD IS MANAGEMENT, ACTUALLY

A note calling AI work leadership drew 283 points and an argument about the word.
The story: Allen Bargi’s note is a two minute read with one move. Code used to deliver certainty. AI does not, so context, clarity and feedback matter more. Bargi rules out anthropomorphism inside the argument itself. The thread took 184 comments to disagree about the noun.
Three camps formed. phil21 said LLMs require zero leadership skills but do require strong management skillsets. simonw countered that people with real management experience get better results from agents. josejux, who shipped a production API this way, said it felt like code review at much higher volume.
The details:
- The failure mode has a body count. boron1006 describes an engineering lead with 25 years of management experience and no coding experience. That produced over 60,000 vibecoded lines in 3 weeks, then a project overrun. One account, uncorroborated.
- Adoption and trust moved in opposite directions. Stack Overflow’s survey put AI use past 84% while trust sat at 29%. Trust fell 11 points in a year.
- The practitioners agree on artifacts, not vocabulary. pixelready describes writing acceptance criteria and wiring up tools the agent can check itself against. alertchecker keeps a Cursor rules file built from what the model got wrong.
- The byline is contested. Several commenters posted AI detector results claiming the note is machine written. The detector pages were not verifiable from the thread.
Why builders care: Solo, there is no second reviewer, so the review harness is the first thing to build. Turn every correction into a checked-in rule, example or test instead of retyping it next session.
TRENDING TODAY
- 🦀 A native app for coding agents - Waku is a desktop app, not another terminal tab. It went up as a Show HN with 18 points and 5 comments. Thin idea, real gap: running three parallel agent sessions in a terminal is genuinely miserable.
- 🧫 An at-home test for infected ticks - You test the tick, not yourself. That is a far easier signal to get. It pulled 231 points and 82 comments, which is a lot for a story with no code in it.
- 📱 A daily digest app for agentic coding news - One push a day, zero to six headlines. Free, no data collection, built to kill the author’s own FOMO. We are contractually obliged to note that the format has legs.
STACK OF THE DAY
📟 claude-trofeo-hud
A live Claude usage HUD that renders to a $38 Thermalright Trofeo Vision LCD. It puts your usage on a dedicated little display instead of another browser tab. It went up as a Show HN with 12 points and 3 comments. That is the usual score for something this niche. Free and open source, and exactly the thing you want the week you start running agents overnight.
Not sponsored. We just feature tools builders would actually use.
BOOKMARKED TODAY
- 🥑 An open-source Grok Bot alternative - Bring your own inference. The builder burned almost 50% of a paid Cursor subscription in 3 hours, so guaca runs on their own keys.
- 👻 A spectre is haunting Unicode - The ghost characters problem, written up properly. It took 195 points and 65 comments, and it is the kind of encoding rabbit hole that eats an afternoon.
- 🌊 Super El Niño is still growing - New forecasts reach record range for the 2026 to 2027 winter. It pulled 123 points and 70 comments, and it is a Q4 planning input if your demand is weather-shaped.
That’s the board for today. Go build something.
Work from any WiFi like it's your home network. NordVPN's Meshnet runs a free private mesh between your laptop, dev box, and home server. SSH from a cafe without exposing a port, the way you'd use Tailscale. The paid VPN on top lets you test geo-fenced Stripe checkouts or feature flags from any country.
We get a cut if you sign up. Only added for tools we use ourselves.
Curated by AI, built by a human.