A Google Team Measured Half of My Argument, and Left the Other Half Open

A Google Team Measured Half of My Argument, and Left the Other Half Open

A recent Google paper showed spec-driven test generation raised bug detection by 9.8 percentage points on a sample from their codebase. But their spec is read out of the code, which leaves the half I care about unresolved. I have been arguing for months that a test suite written by the same model that wrote the code cannot really disagree with it and to support my claim I've developed a new open source Python library.First, a little background. Software testing has been moving in one direction for twenty years, and Specification-Driven Development (SDD) is where that movement has recently arrived. This article argues for one more step: Independent SDD, or ISDD.TDD (Test‑Driven Development) said the tests are the specification. Write them first and the design follows. It worked, and it left the specification in a form only a programmer could read, which meant the person who knew what the software was for could not check it.BDD (Behaviour‑Driven Development) was the answer to that. Structured, plain language scenarios written in Gherkin, using "Given, When, Then" to describe system behaviour. Business analysts, developers and testers argue over the same artifact before any code is written, and what they agree on is the expected behaviour.SDD (Specification‑Driven Development) is the version that arrived with the agents. A formal specification, often in EARS (Easy Approach to Requirements Syntax) notation. It is precise enough that a machine can plan from it, break it into tasks, generate the tests and generate the code. The spec stops being just a formality beside the work and becomes the thing the work is actually generated from.What's interesting to note here is that each step moved the source of truth further upstream: from the tests, to a shared description, and finally to a machine-readable contract.This is the right direction, but my argument is with what happens in SDD using typical AI workflows. Currently, Spec-Driven Development divides the work. It does not divide who has the knowledge. The same specification goes to the planner doing the task breakdown, the test generator and the coding agent. That's "divide and conquer" applied to the workflow pipeline. The knowledge base is left whole, and handed to everyone.Divide and conquer only works when the line you cut along is the line the failure runs across. Here the failure is that the code and the tests come from one reading of the same ambiguous sentence, hence the line runs across knowledge, not work. My argument is to cut along that line: give the coding agent the decisions (the requirements) and withhold the consequences (the acceptance criteria), and the test suite becomes something that can tell the code it is wrong.One word needs pinning down before I get to the paper. A contract is a statement of what the code must do. Google's agent writes its contract by reading the implementation. My contract is written before the implementation exists, and never derived from it. Same word, opposite provenance, and the provenance is the entire argument.My contract also arrives in two halves. The requirements are the decisions somebody made, and the coding agent reads them. The acceptance criteria are the consequences that follow, and it never sees them.A few weeks ago, a team at Google measured the step before withholding: whether writing the contract down helps at all. "Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation" (arXiv 2608.17177) does something narrower than my argument, and it measures what it does. Instead of prompting an agent to write tests, they first ask it to reason about the code and write down its contract: pre-conditions, post-conditions, and the behaviours that are simply undefined. That document becomes what they call a cognitive scaffold and the tests are generated from it.The results, on production bugs from Google's own codebaseThat last number means that more than half the time, a test suite generated from a written contract beat the tests a person wrote.Why I am pleased about results I did not produceThe paper tests the half of my claim I have not been able to test yet. Here is what I have been saying, in two parts:First: a test written from a stated specification is better than a test written from an impression of the code. This is what Google measured, and a specification written down is what they call a contract.Second: the contract has to be hidden from whoever writes the code (otherwise the two agree by construction and the test cannot fail).The first is now measured, by a serious team with a real bug corpus, real statistics and no stake in my conclusion. It came out in the direction I claimed, with a p-value. My own evidence for it was limited to internal runs on my own tasks, which is why somebody else's corpus matters.Why it does not settle the thing I really care aboutLook at where their contract comes from. The agent reads the source and documents what it finds. Pre-conditions, post-conditions, undefined behaviour, all of it derived from the implementation in front of it.So the scaffold is real and it works, but it is downstream of the code implementation. If the implementation resolved an ambiguity the wrong way, a contract written by reading that implementation records the wrong resolution. The test then confidently passes, and as I explained in Part 1 of this series, the green test suite means nothing since the loop never saw a statement of correct behaviour that was independent of the thing being judged.Google's own framing is honest about this. The scaffold is described as improving the agent's reasoning, not as an independent standard. That is a claim about attention, and it is a good one. Independence is a different claim, and the paper does not make it.Nobody has measured whether withholding the criteria from the code-writing agent catches more real defects than a test suite written with full sight of the code. I have measured the neighbouring question, whether refining the criteria sharpens the test suite, across several rounds of planted-fault experiments. I have not found clear evidence of an effect yet. I still expect one to be there, and more than once an early run looked like it, but each apparent gain disappeared when I compared the refined-criteria runs and the original-criteria runs on matched artifacts.I have since measured something adjacent, and it is worth separating from what I have not measured. Across twenty runs on two tasks, one agent wrote the implementation from the requirements alone, and a second agent wrote the test suite from the acceptance criteria, which the first never saw. The suite disagreed with the code every single time, ten out of ten on each. By disagreed I mean a test failed: the suite asserted one behaviour and the implementation did another. In each case the coding agent had also recorded the judgement call in advance, which is what makes the disagreement meaningful: the agent found the ambiguity itself, resolved it, and had no way to check the resolution. The suite was the check.The two tasks disagreed in different ways. On one it was a plain defect: banker's rounding where money needs half up. On the other it was no defect at all, but a decision nobody had made, sitting unnoticed in the difference between < and <=.That is a smaller claim than the one I would like. It says the suite finds disagreements between the implementation and the specification, and that in both of the cases I looked at, the disagreement was worth having: one a defect, one a decision nobody had made. It does not say the suite catches more bugs than a suite written the ordinary way, which is still the open question and still nobody's result.One shortcoming I noticed during these experiments is that a criterion is only as good as the data that exercises it. A criterion about scientific notation cannot fire on data containing none, and a test written from that criterion will pass whatever the code does. That led me to add a new flag to qikly, --propose-fixtures, which finds the criteria your data cannot reach.My takeaways from the paperThree things, and the third one is key.Writing the specification down beats leaving it implicit. That is now measured rather than asserted, by somebody else, and it is a better citation than anything I could produce about my own tool.Where the specification comes from is a separate question, and Google's result does not touch it. Their contract is read off the code, which is the cheapest source and the one source that can never tell you the code is wrong.And who is allowed to read it is a third. TDD, BDD and SDD each moved the specification further upstream, and none of them had to answer this, because a person who writes both the specification and the code still meets code review, a tester, and a colleague who reads the same sentence differently. An agent meets none of them. Divide the work all you like; the division that decides whether green means anything is the one across what each agent knows.Getting scooped on half my argument was actually a good day. I spent an afternoon deciding whether this paper undercut what I have been building. Turns out that it actually does the opposite. It removes a claim I was making based on limited research and replaces it with one somebody measured, which leaves me arguing for one thing instead of two. A smaller claim that is still standing is worth more than a large one nobody has tested.Independent Spec-Driven Development (ISDD) In summary, the evolution I am proposing is Independent Spec-Driven Development, ISDD: the same specification, divided so that whoever writes the code cannot read the half that judges it.Whether that division catches more real defects is the open question, and it is what qikly is built to test. If you know of work that measures it, I would be keen to read it.Previous articles in this seriesPart 1 explains my argument in detail and claims "the green test suite means nothing":https://towardsdatascience.com/towards-spec-driven-test-automation-part-1/Part 2 walks through a full run of qikly: https://towardsdatascience.com/towards-spec-driven-test-automation-part-2/Gal Arav is the author of Applied Statistics for Data Science and maintains qikly, a new open-source Python project for spec-driven test automation, at: https://test.qikly.com/?ref=tds_spec_part3

Original Source

Read the full article at Towardsdatascience →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.