1 comments

  • jkalichman 2 hours ago
    A bit more context on why we built this:

    A lot of the recent signals point in the same direction: back in Dec 2025, Boris Cherny said 100% of his Claude Code contributions "over the previous 30 days" were written by Claude, and as Garry Tan recently mentioned - YC has already seen companies with 95% AI-generated codebases.

    We’re getting very good at the "how".

    But most of the systems around software development still optimize for shipping faster and more reliably. They don’t really answer the more basic question: was this worth building?

    That’s the idea behind what we call the Evidence Loop:

    real users (or synthetic copy of it) → evidence → spec → agents build → observe what happens → start the loop again.

    We’re also researching Synthetic Twins. We’ve published peer-reviewed work (currently in press). I don’t think synthetics should replace real users; the interesting part is using them to explore hypotheses cheaply, then bringing those hypotheses back to real users and real behaviour.

    Longer term, the bet is simple: agents shouldn’t just get a ticket saying “build X.” They should have access to the evidence explaining why X should exist at all.

    Thank you!

    Julian

    DM's are welcomed: https://www.linkedin.com/in/j16h/