#123

Anthropic's AI group lost 17 points, and OpenAI's ten proofs ship no failure count

Anthropic's AI group scored 50% on a quiz about code it just wrote, against 67% for hand coders. OpenAI shipped Lean files for ten proofs and no failure count.

Listen to this edition

Two engineers can run the same assistant and look identical from the outside. On the same quiz they land 17 points apart. Anthropic measured it on 52 mostly junior engineers: the AI group averaged 50%, the hand coders 67%.

The post that put this on the front page today is Sean Goedecke’s. The framing there is blunter: “the human is the bottleneck, not the model.”

In today’s indie hacker news:

  • 🧠 The AI group lost 17 points on mastery
  • 🔓 A nightly cron that carries your fork
  • 🧮 Ten proofs, an invoice, and no failure count
  • 🧵 289 factories will sample, only 189 will make
  • 💪 A trainer priced the app against Netflix
  • 🔍 Ask your Neovim config what it does

TOP STORIES

SKILL ISSUE, LITERALLY

🧠 Four studies now price what domain expertise is worth against an LLM

Four studies now price what domain expertise is worth against an LLM

The story: LLMs reward expertise went up on July 24 and hit the front page today. Goedecke’s argument is that models turn everybody into a generalist, so the scarce input is what you already know. The illustration is Terence Tao’s ChatGPT conversation about a counterexample to the Jacobian Conjecture. Goedecke could not reproduce it, with unlimited tokens to burn.

The techniques in that transcript are all structural. Tao sends short messages that answer the gist rather than each point. Pushback arrives as “this looks more complex than I was hoping for,” never as a flat contradiction. And Tao almost never takes the model’s advice about where to go next.

“By signalling expertise, Tao shunts the model into ‘talking-to-mathematicians’ mode, not ‘explaining-to-amateurs’ mode.” Sean Goedecke, seangoedecke.com

The details:

  • The habit that split the field: clusters that delegated averaged under 40% on Anthropic’s quiz. Clusters that asked conceptual or follow up questions scored 65% or higher.
  • The one variable: one high scoring cluster differed from the worst only in asking follow up questions afterward.
  • Where the gap was widest: debugging, which the study flags as its particular concern. Debugging catches generated code being wrong.
  • What the speed was worth: about two minutes, and the difference did not reach statistical significance. The study sizes the mastery loss at nearly two letter grades.
  • The counterweight: METR clocked 16 maintainers 19% slower with AI across 246 issues in repos they knew cold. It has since marked that result out of date.

Why builders care: The diff is the same either way, so nothing in your toolchain can catch the expensive pattern. Stanford’s payroll data says the entry level jobs where that skill used to get built are thinning out first.

Somebody’s fix for this is in Trending today, and it is to retype the model’s code by hand.

FORK IT AND CRON IT

🔓 The case for open devtools now fits inside one nightly cron prompt

The case for open devtools now fits inside one nightly cron prompt

The story: Devtools must be open source published August 2, and the Hacker News thread pulled 189 comments. David Crawshaw’s claim is that agents killed both the upfront and the ongoing cost of personalizing software. Shipping source is now the only way an end user product makes sense. Crawshaw names the wall: Claude Code is closed, so you do not get to personalize it.

The whole mechanic is two prompts. Build the thing from source and record why you changed it. Then set a nightly cron to fetch upstream and rebase your local changes on top. The proof is Shelley, Apache licensed at 590 stars and 667 releases auto cut on every commit to main. Its README admits much of it was written by Claude Code and Codex.

The details:

  • The prompt doing the work: “fetch upstream changes to the <software> and rebase all local changes on top.” The cron then verifies the build and swaps it in.
  • The working recipe: Fabricio20 carries roughly 6 forks on Stacked Git, and says Claude reapplies the patch stack. A total break costs under an hour.
  • The other side, stated in public: trjordan is keeping tern.sh closed so it stays hosted and team friendly. dregitsky says boxes.dev ships too fast for anyone customizing it to keep up.
  • What closed costs: dipanshuhappy’s employer discouraged Claude Code specifically because it has not been open sourced.
  • The live warning: Google’s Gemini CLI repo stays Apache 2.0 while the consumer tiers move to a Go successor. Whether that successor will be open source is the top reply at 50 votes, unanswered.

Why builders care: Every Apache licensed dependency in your stack is a fork you can carry now. The feature you have been waiting on a maintainer to merge is a cron job instead.

simonw argued the other direction in the same thread. Publish a clever indexing scheme, and a rival’s agent imitates it in minutes.

SHOW YOUR WORK, NOT YOUR MISSES

🧮 OpenAI shipped ten Lean files anyone can rebuild, and no failure count

OpenAI shipped ten Lean files anyone can rebuild, and no failure count

The story: Ten advances in mathematics and theoretical computer science went up August 1. The credited system is an internal version of Astra, a model OpenAI has not released. Each of the ten, OpenAI says, resolves or makes substantial progress on a long standing open problem. The cost line, verbatim: the tokens needed to solve these problems would run roughly $2,000 at Sol API rates.

The part builders can copy is the certificate. openai/ten-proofs ships one Lean 4 file per result under Apache-2.0. A skeptic runs the build instead of trusting a blog post.

The details:

  • The exact toolchain: Lean 4.32.0 plus mathlib and Lake rebuilds all ten formalizations. A separate directory covers an independent prover.
  • What stayed unpublished: the attempt count, the success rate and the prompts. Simon Willison asked how many problems got the same spend without reaching a solution.
  • The figure is contested: OpenAI attaches roughly $2,000 to the whole set. Willison reads it as under $2,000 on each one, a factor of ten apart.
  • A repo one commit deep: 442 stars a day after publication, off exactly 1 commit from 1 contributor.
  • The competitive frame: Anthropic reported burning $100,000 in tokens on cryptography work days earlier. That is roughly 50 times OpenAI’s figure for all ten.

Why builders care: Shipping a machine checkable artifact next to your claim works on types, tests, schemas and migrations. The denominator is the expensive half, and nobody publishes that unprompted.

289 WILL SAMPLE, 189 WILL SHIP

🧵 A fashion grad asked r/startups where to start and got trunk shows

A fashion grad asked r/startups where to start and got trunk shows

The story: A recent fashion school graduate asked r/startups where to begin with a wholesale brand. They have a name, social accounts, a stated taste for denim and heavy embellishment, and no factory. The target is people aged 17 to 35 who lean alternative, out of rural Georgia. The first reply told them they were in the wrong subreddit. The only product advice in the thread was local marketing. Trunk sales in hip areas, trades with coffee shops, bands and local influencers.

Nobody in the thread mentioned capacity, the label, or what heavy embellishment costs you. The CFDA production directory is open access and lists over 380 U.S. contract manufacturers. The FTC wants fiber content, country of origin, and either a legal company name or an RN. An RN is a free identifier you print instead of the name, and some buyers require one.

The details:

  • Where the funnel narrows: 289 CFDA listings will cut a sample. Only 189 accept a 10 to 50 unit production run.
  • The only two filters: the directory sorts by New York at 260 listings and Los Angeles at 83. There is no Southeast option, and the poster does not want to drive far.
  • The thinnest pools are the aesthetic: 65 listings offer denim services and 135 do embellishment. Sample making runs 277.
  • Mislabeling is priced per garment: an FTC order violation runs up to $53,088. Each mislabeled piece counts on its own, so a bad 50 unit drop is 50 violations.
  • The signature is the liability: decoration escapes fiber disclosure below 15% of surface area and 5% of fiber weight. Describing it in ad copy triggers the disclosure anyway.

Why builders care: In apparel the label and the FTC paperwork gate the wholesale account. None of it has a free tier, and it all lands before your first order.



FIRST DOLLAR

PRICED AGAINST NETFLIX

💪 A former trainer rebuilt the job as an app and undercut the old rate

FitBySci comes from a former personal trainer who left the field for unrelated work. Then came a kid at 40 and the realization that a decade had quietly added about 75 lbs. After getting back in shape, they built the app to coach other people at a fraction of the old price. It is priced one cent above the base Netflix subscription, and the first members came from their own circle.


STACK OF THE DAY

🔍 Ask your own Neovim config what it does

how.nvim answers questions about your Neovim setup. How a thing works, which keymaps exist, what a small config change would look like. It is the author’s first Neovim plugin. The pitch is the moment you are deep in work and will not stop to read docs.

Not sponsored. We just feature tools builders would actually use.


BOOKMARKED TODAY


See you tomorrow. Hit reply and tell me what you shipped this week.

Curated by AI, built by a human.