Skip to content
semi-sentient code

How I Trust Code I Didn't Write

· 17 min read

#testing #ai-engineering #python

A while ago I built a small toy project for running agents on Temporal and wrote about it on my old Substack. It’s a deep-research agent whose workflow fans a question out to planning, research, validation and critique agents, then stitches the answers back together. A lot has changed since then, especially the quality of coding harnesses and frontier models.

When I wrote that post, I was still deep in the code. I spent hours reviewing it, and built systems to refactor it and condense what I’d learned. I don’t work that way anymore. I still read code, but far less of it line by line, because with AI development moving this fast there’s simply too much to keep up with. I’ve found myself caring less about how the code works and more about two questions: does it do what I need, and will this change break something I didn’t expect?

To me, that’s the main problem with AI coding right now. How do we define what correct means for our system? And how do we put that definition into the harness, with a feedback loop, so the agent stays in its lane?

I believe in a Swiss cheese approach. No single layer of defense catches every bug, but enough layers stacked together catch most of them. Today a lot of developers lean on TDD, integration tests, agentic code review and agentic loops to land a feature. All of that is great. But I’ve also found real value in a few testing techniques from the old days that make me more confident in what I ship. The biggest one starts with a question: what if, instead of testing examples, you tested the properties of a system?

That’s what this post is about. In my opinion, testing properties is the best way to pin down what correctness means for a system. Everything below runs, and the code is linked at the end.

The weak layer

When I review AI-written code I think in layers, bottom to top:

  1. Types and schemas. Pydantic models, type hints. They catch the shape of a mistake.
  2. Example tests. assert f(2) == 4. The AI writes these for free.
  3. Property tests. For all inputs, something holds.
  4. Mutation testing. Deliberately break the code and check the tests notice.
  5. Stateful tests. Random sequences of operations against invariants.
  6. Human review. Me, reading.

Layer 2 used to be the backbone because it’s the easiest to write. Write a few tests, validate the happy path. If I’m clever, I’ll think of a few unhappy paths too, and over time the suite accumulates the edge cases I missed the first time around. A lot of companies gate merges on these tests, usually through a coverage number. Whether that improves code quality or just adds toil is an open question.

When coding agents started producing mountains of code, the natural reaction was: aha, the agent can write its own tests first, then the implementation. Red, green, refactor, and code quality goes up. That’s partly true. But every problem unit tests always had is still there, and AI adds a few new ones.

First, if the agent misunderstands the intent of the codebase or the feature, the tests just encode that misunderstanding. The tests and the code come out of the same model, in the same sitting, from the same understanding of the problem. If that understanding is wrong, both are wrong in exactly the same way. Everything passes, and the green build tells you very little.

Second, agents can probably write more edge cases than I ever could, but in my experience they gravitate toward the obvious failures: the empty list, the None, the negative number. The subtle cases, the ones that depend on how two parts of the system interact, are the ones that slip through.

Third, nothing stops the agent from changing a test when it gets in the way. This has gotten better over time, but nothing really prevents it, just as nothing prevented a tired engineer from doing the same thing before. Fixtures and immutable test sets help. They don’t fix the context problem, though, and they say nothing about the properties of the system.

What I actually want is different in kind. I don’t want to say “this input gives this output.” I want to say “no matter what you feed this function, this must remain true.” That’s a statement about the problem, not about how it’s solved. I don’t want the agent to get lucky and find a narrow set of examples it can pass. I want the code tested against the inputs the agent didn’t think of.

Property tests aren’t immune to the same failure modes, to be clear. An agent can loosen a property as easily as it can rewrite an example, and a property written from a wrong understanding is still wrong. The difference is that properties are short and about the problem, so a wrong one is easy for a person to spot. That’s why I write or review them myself, and why layer 4 exists.

Layers 3 through 5 are how you get the signal back. Layer 6 is still there; it’s just last, not first, and it’s not enough on its own.

What a property test is

An example test pins down one input and one expected output. A property test states a rule that should hold for every valid input, then uses a library to generate inputs that try to break it. It has three parts:

  1. A generator that describes what valid inputs look like: lists of strings, integers between 1 and 10, a Pydantic model with these fields.
  2. A property that must be true of the output, whatever the input was.
  3. Shrinking. When the library finds an input that breaks the property, it doesn’t hand you that input. It keeps simplifying it until it has the smallest one that still fails.

In Python that library is Hypothesis. In TypeScript it’s fast-check, which works the same way, shrinking included. Everything in this post translates.

The toy

Start small so the mechanics are clear. Here’s a function that removes duplicates from a list while keeping the first occurrence of each item in place:

def dedupe_keep_order(items: list[str]) -> list[str]:
    return list(set(items))

There’s a bug: sets don’t preserve order. It’s a simple thing and easy to overlook.

Here’s the example test most of us would write:

def test_dedupe():
    assert dedupe_keep_order(["a", "a", "b"]) == ["a", "b"]

It looks reasonable, and it might even pass. Python randomizes string hashing per process, so the order a set iterates in can change from one run to the next. This test can be green on your laptop and red in CI, or the other way around. That’s the trap with example tests: they’re simple to write, and it’s just as simple to write one that happens to work, unless you sit down and think through every way the function could fail.

I don’t want to enumerate inputs. I want to say what the function is for, and let the machine go looking for inputs I didn’t think of. For a dedupe, that comes down to three properties:

  • Nothing is dropped. Every item in the input shows up in the output.
  • Everything is unique. No item shows up twice.
  • Order is preserved. Items appear in the order they were first seen.

In Hypothesis:

from hypothesis import given, strategies as st


@given(st.lists(st.text()))
def test_dedupe_loses_nothing(items):
    assert set(dedupe_keep_order(items)) == set(items)


@given(st.lists(st.text()))
def test_dedupe_has_no_duplicates(items):
    out = dedupe_keep_order(items)
    assert len(out) == len(set(out))


@given(st.lists(st.text()))
def test_dedupe_keeps_first_seen_order(items):
    out = dedupe_keep_order(items)
    first_seen = []
    for x in items:
        if x not in first_seen:
            first_seen.append(x)
    assert out == first_seen

@given(st.lists(st.text())) is the generator. It tells Hypothesis to build lists of strings, hundreds of them, and run each test against every one. The first two properties pass. The third does not:

    @given(st.lists(st.text()))
    def test_dedupe_keeps_first_seen_order(items):
        ...
>       assert out == first_seen
E       AssertionError: assert ['', '0'] == ['0', '']
E       Failing test case: test_dedupe_keeps_first_seen_order(
E           items=['0', ''],
E       )

The failing input is ['0', '']: two short strings that the function returns in the wrong order. Hypothesis found a failure, then shrank it, trying smaller and simpler inputs until it reached the smallest case that still fails. That minimal example is what makes property tests practical to debug. It’s an exact reproduction, and it’s what I hand back to the agent to fix.

The fix is one line, list(dict.fromkeys(items)). Once the three properties exist, any future rewrite, by me or by an agent, has to satisfy them.

The same bug, in real code

The toy was dedupe_keep_order. The project has the same function in disguise. A research run produces a list of steps, each with its own sources, and the final report needs one de-duplicated list of citations across all of them:

class ResearchResult(BaseModel):
    ...
    @property
    def all_sources(self) -> list[Source]:
        """Aggregate all unique sources from all research steps."""
        seen_urls: set[str] = set()
        unique_sources: list[Source] = []
        for step in self.steps:
            for source in step.sources:
                if source.url not in seen_urls:
                    seen_urls.add(source.url)
                    unique_sources.append(source)
        return unique_sources

Claude wrote this. It’s correct, but there were no tests for it.

The properties are the same three as the toy, plus one I’ll explain in a moment. The only new work is telling Hypothesis how to build a ResearchResult:

# A small URL alphabet makes duplicates across steps likely,
# which is the interesting case for de-duplication.
urls = st.sampled_from([f"https://example.com/{i}" for i in range(8)])

sources = st.builds(Source, url=urls, title=st.text(max_size=20),
                    description=st.text(max_size=40))

steps = st.builds(ResearchStep,
                  iteration=st.integers(min_value=1, max_value=20),
                  question=st.text(min_size=1, max_size=50),
                  findings=st.text(max_size=100),
                  sources=st.lists(sources, max_size=5),
                  confidence=st.floats(min_value=0.0, max_value=1.0))

results = st.builds(ResearchResult,
                    query=st.text(min_size=1, max_size=50),
                    steps=st.lists(steps, max_size=6))

st.builds takes a class and a strategy per field; anything you leave out is either inferred from its type hint or left at its default. The comment on urls matters: if you let Hypothesis generate arbitrary URLs, it will almost never produce two the same, and your de-duplication code will never actually de-duplicate anything. Choosing a tiny alphabet is how you steer it toward the collisions you care about. A lot of the work in property testing is picking generators that make the interesting cases common.

Then the properties:

@given(results)
def test_all_sources_loses_nothing(result):
    expected = {s.url for step in result.steps for s in step.sources}
    assert {s.url for s in result.all_sources} == expected


@given(results)
def test_all_sources_has_no_duplicate_urls(result):
    urls_seen = [s.url for s in result.all_sources]
    assert len(urls_seen) == len(set(urls_seen))


@given(results)
def test_all_sources_keeps_first_seen_order(result):
    first_seen = []
    for step in result.steps:
        for s in step.sources:
            if s.url not in first_seen:
                first_seen.append(s.url)
    assert [s.url for s in result.all_sources] == first_seen


@given(results)
def test_all_sources_is_idempotent(result):
    """Feeding the de-duplicated list back in as one step changes nothing."""
    once = result.all_sources
    one_step = ResearchStep(iteration=1, question="q", findings="",
                            confidence=1.0, sources=once)
    again = result.model_copy(update={"steps": [one_step]}).all_sources
    assert again == once

Idempotence is the one I’d recommend adding to any “clean up a collection” function. It says: once this has run, running it again is a no-op. It’s cheap to state, and as you’ll see below, it’s one of the properties that catches the planted bug.

Breaking it on purpose

The code was correct, so to show the tests working I had Claude plant a bug. It moved one line:

        unique_sources: list[Source] = []
        for step in self.steps:
            seen_urls: set[str] = set()      # <- was above the loop
            for source in step.sources:

That’s a believable refactor slip. Sources are now de-duplicated within a step but not across steps. The example test you’d naturally write still passes:

def test_all_sources_dedupes():
    src = Source(url="https://example.com/a", title="a", description="")
    step = ResearchStep(iteration=1, question="q", findings="",
                        confidence=1.0, sources=[src, src])
    result = ResearchResult(query="q", steps=[step])
    assert [s.url for s in result.all_sources] == ["https://example.com/a"]

One step, a repeated URL, green. The bug only shows up when two different steps share a source, and that’s not a case most people think to write by hand.

Three of the four properties fail. The one that passes is “loses nothing,” and that’s expected: the bug adds duplicates, it doesn’t drop anything. Each property guards against a different kind of failure, and no single one catches everything. That’s the Swiss cheese idea again, at the scale of one function.

Here’s the shrunk case from the no-duplicates property:

E       Failing test case: test_all_sources_has_no_duplicate_urls(
E           # The test always failed when commented parts were varied together.
E           result=ResearchResult(
E               query='0',  # or any other generated value
E               steps=[ResearchStep(
E                    iteration=1,  # or any other generated value
E                    question='0',  # or any other generated value
E                    findings='',  # or any other generated value
E                    sources=[Source(
E                         url='https://example.com/0',
E                         title='',  # or any other generated value
E                         description='',  # or any other generated value
E                     )],
E                    confidence=0.0,
E                ), ResearchStep(
E                    ...
E                    sources=[Source(
E                         url='https://example.com/0',
E                         ...
E                     )],
E                )],
E           ),
E       )

Two steps. One source each. Same URL. And read the annotations: # or any other generated value on every field that doesn’t matter. Hypothesis is telling you the exact shape of the bug: two steps sharing a URL, and nothing else.

The test that found a design smell

Not everything Hypothesis finds is a bug. Sometimes it’s a decision nobody made on purpose.

A ResearchQuery has a max_iterations (default 3) and a depth of quick, standard or deep. There’s a validator that bumps the iteration count to match the depth, but only if the user didn’t set it explicitly:

@model_validator(mode="after")
def adjust_iterations_by_depth(self) -> "ResearchQuery":
    depth_defaults = {"quick": 1, "standard": 3, "deep": 5}
    # Only adjust if using default value
    if self.max_iterations == 3 and self.depth != "standard":
        object.__setattr__(self, "max_iterations", depth_defaults[self.depth])
    return self

The property is simple: if I set max_iterations explicitly, it should be respected.

@given(st.integers(min_value=1, max_value=10), depths)
def test_explicit_iterations_are_respected(max_iterations, depth):
    q = ResearchQuery(query="q", max_iterations=max_iterations, depth=depth)
    assert q.max_iterations == max_iterations

It fails right away, on max_iterations=3, depth="deep", and you can see why from the source. The validator can’t tell “the user asked for 3” from “the user didn’t say.” It uses the value 3 as a sentinel for unset. A user who explicitly wants three deep iterations gets five.

Is that a bug? Arguably it’s a reasonable heuristic. But it’s a decision that was never made; the agent reached for the simplest thing that satisfied the docstring, and the docstring was ambiguous. The real problem is that “unset” wasn’t representable, so the code borrowed a real value to mean it.

The fix is to make the absence a real state:

max_iterations: int | None = Field(default=None, ge=1, le=10, ...)

@model_validator(mode="after")
def adjust_iterations_by_depth(self) -> "ResearchQuery":
    depth_defaults = {"quick": 1, "standard": 3, "deep": 5}
    if self.max_iterations is None:
        object.__setattr__(self, "max_iterations", depth_defaults[self.depth])
    return self

Now an explicit 3 is a 3, the API schema defaults to None so depth still drives the default for callers who don’t care, and the property passes for every value. I added its mirror image too, “unset follows depth,” so both halves of the contract are pinned:

@given(depths)
def test_unset_iterations_follow_depth(depth):
    q = ResearchQuery(query="q", depth=depth)
    assert q.iterations == {"quick": 1, "standard": 3, "deep": 5}[depth]

That’s a two-line change to the model and a one-line change to the API, and I wouldn’t have found it without the test, because the old behaviour looked fine in every example.

Property testing doesn’t only find bugs. It also surfaces decisions that were never made explicitly, so a person can make them.

Do your tests actually catch anything?

Property tests are my layer 3. Layer 4 asks a different question: if the code were wrong, would these tests notice?

Mutation testing answers it by brute force. A tool takes your source, makes one small change (flip a < to <=, replace a constant, swap and for or), runs the tests, and records whether anything failed. If a test fails, the mutant is “killed.” If everything stays green, it “survived,” and you’ve found code your tests don’t constrain. Repeat for every mutation the tool can think of.

I ran mutmut over models.py, the file with all the pure logic in this project. First with the only tests that existed for it, five example-based tests the agent had written, and then with the property tests added.

Test suite Mutants killed Survived Never reached
Example tests (5, AI-written) 20 / 44 0 24
Property tests (first pass) 39 / 44 4 1
Property tests (final) 43 / 44 1 0

The first row is the “layer 2 is weak” argument in one line. Twenty-four mutants were never even reached: whole functions with no test touching them. The example tests looked fine. Each one tested a real behaviour with a real assertion. They just covered a small, hand-picked slice of the file and left the rest untouched. That’s the failure mode of AI-written tests: not that they’re wrong, but that they’re narrow in a way that’s invisible until you measure it.

The middle row is the interesting one. My first pass of property tests left four survivors, all on the same line:

-        citation = f"[{i}] {source.title or 'Untitled'}"
+        citation = f"[{i}] {source.title or 'XXUntitledXX'}"

I’d written a property that the citation numbering was [1], [2], [3]… and never one that said the title had to be in there. Mutation testing caught a gap in a test I had just written. That’s the second argument for layer 4: it audits layer 3 too, and property tests are not immune to being too loose.

The last row has one survivor, and it’s worth looking at:

     def completion_percentage(self) -> float:
         if not self.sub_tasks:
-            return 0.0
+            return 1.0

No test can kill this, because no test can reach it. sub_tasks is declared with min_length=1 in the Pydantic model, so a plan with zero sub-tasks can’t be constructed. The guard is dead code. That’s called an equivalent mutant, and when you hit one you have two options: delete the dead branch, or leave the mutant surviving and write down why. I left it, because deleting the guard makes the function look unsafe to a reader who doesn’t know the schema. A 43/44 with a reason beats a 44/44 bought by removing a defensive line.

One caveat on the numbers: mutmut skips decorated functions, so all_sources (a @property) and the validator above aren’t in the 44. They’re tested; they’re just not audited by this tool. Know what your score counts.

Testing the thing that changes over time

The models so far are value objects: build one, ask it a question, done. A Temporal workflow isn’t like that. A ResearchPlan lives for minutes or hours while tasks get added, completed, and queried. The bugs there are sequence bugs: “this is wrong after you do A, then B, then A again.”

Hypothesis has a tool for that, and it’s my layer 5. You describe the operations as rules and the things that must always be true as invariants, and it generates random sequences of operations, checking the invariants after every step:

from hypothesis.stateful import RuleBasedStateMachine, invariant, precondition, rule


class PlanMachine(RuleBasedStateMachine):
    def __init__(self):
        super().__init__()
        self.plan = ResearchPlan(original_query="q", query_analysis="a",
                                 sub_tasks=[ResearchSubTask(task_id="t0", description="seed")])

    @rule(deps=st.lists(st.integers(min_value=0, max_value=5), max_size=2))
    def add_task(self, deps):
        existing = [t.task_id for t in self.plan.sub_tasks]
        self.plan.sub_tasks.append(ResearchSubTask(
            task_id=f"t{len(existing)}", description="d",
            dependencies=[existing[i % len(existing)] for i in deps]))

    @precondition(lambda self: self.plan.get_ready_tasks())
    @rule(data=st.data())
    def complete_ready_task(self, data):
        task = data.draw(st.sampled_from(self.plan.get_ready_tasks()))
        task.is_completed = True

    @invariant()
    def ready_is_subset_of_pending(self):
        pending = {id(t) for t in self.plan.get_pending_tasks()}
        assert all(id(t) in pending for t in self.plan.get_ready_tasks())

    @invariant()
    def ready_tasks_have_completed_dependencies(self):
        done = {t.task_id for t in self.plan.sub_tasks if t.is_completed}
        for task in self.plan.get_ready_tasks():
            assert set(task.dependencies) <= done

    @invariant()
    def percentage_is_consistent(self):
        pct = self.plan.completion_percentage()
        assert 0.0 <= pct <= 100.0
        assert (pct == 100.0) == (not self.plan.get_pending_tasks())


TestPlanMachine = PlanMachine.TestCase

About forty lines. Two operations, three invariants. Hypothesis will add tasks with random dependency graphs, complete them in random valid orders, and after every step confirm that “ready” tasks are a subset of “pending” ones, that nothing is ready while its dependencies are open, and that the percentage hits 100 exactly when nothing is left. If it ever finds a sequence that breaks one, it shrinks that too: down to the shortest sequence of operations that reproduces it.

This one passes. What I get from it is an executable description of how a research plan should behave. It took about as long to write as three example tests, and it still applies if the implementation changes.

Fuzzing a calculator that uses eval

Fuzzing is closely related to property testing, with less structure: throw random input at code and see if it crashes. There’s a calculator tool in this project that the agents can call, and it is implemented with eval, in a restricted namespace, wrapped in a try/except that promises to never raise. That promise is worth testing:

@pytest.mark.asyncio
@settings(max_examples=300, deadline=None)
@given(st.text(max_size=40))
async def test_calculate_never_raises(expression):
    out = await calculate(expression)
    assert out.startswith(("Result: ", "Calculation failed: "))

Three hundred random strings, including ones that look like Python, and it never raised. That shows it doesn’t crash, not that it’s safe; a test like this says nothing about whether someone could escape the restricted namespace. I’m still not comfortable with eval being in there, but a test that checks the promise is better than reading the code once and deciding it looked fine.

What this changes about working with an agent

Here’s the loop I actually run now.

I describe the feature. The agent writes the code and some example tests. I don’t read the example tests closely; I treat them as smoke. Then I write the properties, or increasingly I ask the agent to propose properties and I edit them, which is a better use of my time than reading every function. Properties are short, they’re about the problem rather than the solution, and a wrong one is obvious in a way a wrong example rarely is.

Then I run the mutation score, because the properties can be too loose and the tool will tell me where.

What I’ve noticed is that the agent gets better at the code when the properties exist first. “Make all_sources satisfy these four invariants” is a better prompt than “aggregate the unique sources.” It’s an executable spec. When the agent gets it wrong, the shrunk counterexample goes straight into the next prompt, which gives a tighter feedback loop than reviewing the code myself.

That’s the main point of this post. Property tests, mutation testing, and stateful tests aren’t only a defense against AI-written code. They’re a way to tell an AI what correct means, in a form it can check its own work against. The agent does most of the writing; these layers are how I verify it.


The code in this post is real and lives in the repo: tests/test_properties.py, tests/test_plan_state_machine.py and tests/test_calculate_fuzz.py, with make mutate to reproduce the table. The planted bug is not committed; everything else is.

← all writing