<- Back
Comments (25)
- kqrThis is an interesting idea. It's effectively what a good software engineer already does in their head, except it's doing it with a real compiler.Richard Gabriel wrote something that has really stuck with me:> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!
- faremintUnsteered coding agents like to write unit tests on the smallest possible functions (based on their training data) - but that was important when humans were writing code. AIs rarely make mistakes on simple functions, what is more important is testing the most outer layers, the business requirements rather than implementation.
- olalondeThis isn't entirely new. I can't be the only one who has built throwaway UIs to test backend APIs instead of writing proper tests...
- glownaggerThis is an application of the rule of three: you need at least three clients to write a good library, protocol, or other underlying abstraction. Except it proposes using a large language model to write the three clients and then throwing them away.
- tartakovskybut what if you’re building a harness… the outer layer involves a human in the loop no?
- folkravUnless I'm missing something, this doesn't help with preventing regressions. In the end, as the author already puts it, it's an integration test in the end, why not just write the integration tests directly?
- simonwI've been using this pattern quite a bit recently for API design, and I really like it.The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
- anilakar> A library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures.Sigh.Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
- bunderbunderI had good results with a similar technique this summer.I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
- CBLTI just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
- bakiesNot unthinkable at all. Been doing this for several years as an opt-in on pr tags usually but now with agents writing code I have it on by default. I've got several skills files about accessing with a read only account and gitops done through the PR. It's an excellent dev environment for the agents.
- exacWe use NX in our monorepo, and it is great at determining which testing/linting tasks need to be run based on which libraries in the repo were "affected". We have a bunch of e2e tests, and I've been having a lot of success getting Claude with Opus to run only the relevant e2e tests when appropriate during development.
- aliasxneoI've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
- MuromecThat's a thing with AI-generated code. The lying machine happily reports that "boss, everything is clean, tests pass, code gate green", but once we need to build on top it's always "preexisting flaky tests, not related to this section" and trying to do commit --no-veriry before it's slapped on it's little robot hands twice.Gets annoying pretty quickly
- skeledrewI'm just actually using the thing I'm building while building it, so I feel the sharp edges and the project quickly evolves based on actual need. Nothing new really.
- stabblesAnother trick I have found useful is to do integration testing with coverage enabled. One agent creates tasks for subagents that run the application with coverage enabled, it merges the reports, and based on that it comes up with new tasks for subagents, repeat until coverage no longer moves.
- hugsi call this (new?) type of test the "vibe check"
- rgoulter> In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.Eh. I think you're just going to end up with slop, or sloppy recommendations?My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
- SA9GSeems like one more arrow in the toolkit. But the best testing looks at the sources and explores the cracks between the strata with edge cases, looks at limits, and where one method changes to another. (And hats off to Murphy, for waiting until after you ship...)
- theodorewilesi'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
- palaceme[flagged]
- haukebri[flagged]
- anonundefined
- anonundefined
- anonundefined
- ardub[flagged]
- folayii[flagged]