Engineering · Testing infrastructure

How 7,400 tests became a ten-minute decision

Korivo did not have a shortage of tests. It had a system that kept paying for the same evidence, rebuilt source when it did not need to, and treated every change as if it might be a release. The fix was not fewer tests. It was deciding which evidence a change actually needed.

Korivo has roughly 7,400 automated tests, plus variable-count end-to-end, UI journey, mutation, property, and soak testing. Most unit tests are quick. The problem was everything around them: cold builds, simulator setup, repeated work, and a 14-scenario end-to-end suite that could take an hour or more when it ran repeatedly to check deterministic results.

I took a little from what I had observed at Meta and other large engineering organizations and asked for a very dumbed-down test selection engine. I gave the direction to the LLM rather than specifying the implementation. There is no probabilistic model deciding what is safe. It is a committed map, four risk levels, a cost budget, a small failure-history signal, and a refusal to pretend when the answer is too expensive.

The test suite was large, but size was not the expensive part

The approximate inventory at the start of this work looked like this. The total moves as features and generated cases change, and the stress suites deliberately have variable counts, so false precision would be misleading.

2,698iOS applicationservices, SwiftData, view models, sync, wire behavior and regressions
4,245StudioTypeScript, React components, stores, codecs and desktop behavior
191Swift engine & policypure kernels, shared fixtures and architectural rules
237Rust & native hostcore invariants, Tauri commands and platform integration
Figure 1. A working snapshot, not a permanent scoreboard. End-to-end repetitions, property cases and soak rounds sit outside these fixed counts.

A useful audit found that about 77 percent of the suite represented unique behavioral pins. The suite was not mostly tax. The waste was concentrated in how it was invoked.

A small Swift test might execute in seconds after the application had spent tens of minutes rebuilding. A fresh worktree often meant a fresh simulator and fresh DerivedData. The same full iOS suite could run several times against unchanged source. End-to-end setup once spent 49 minutes building before scenario one began; another run spent 12 minutes waiting for a frontend that could be built in seconds. The test was not slow. The path to the test was slow.

Cold application build25–60 min observed
Repeated E2E setuppaid per run
Simulator lifecycleclone, boot, teardown
The selected testoften seconds
Figure 2. The proportions are illustrative; the measured incidents are not. Optimizing assertions would have attacked the smallest bar.

The wrong default was “run everything”

Running everything feels conservative. In practice it made feedback so expensive that a routine edit could inherit the cost of every platform, every simulator, and every multi-device journey. The repeated end-to-end runs were especially punishing: 14 scenarios exercise the real iOS and Studio stacks through private folders and simulated devices, and repeated runs check that the same inputs converge to the same result. That is valuable evidence. It is not useful evidence for every CSS change or isolated engine test.

There was also a less obvious correctness problem. Heavy gates competed for the same machine. Concurrent builds produced misleading failures; automated apps could collide with a developer's live Studio state; a stale binary from one build directory could be installed after another directory had been tested. More testing was sometimes producing less trustworthy evidence.

The first fixes therefore attacked mechanics rather than selection. Build once, then run many suites without rebuilding. Reuse DerivedData only when its source and native-framework fingerprint matches. Give every worktree its own simulator clone and private application data. Serialize machine-heavy gates. Record full logs and result bundles while keeping the console terse. Treat a command that executes zero tests as a failure, because “green” with no test cases is not evidence.

A deliberately simple impact map

Test selection works by looking at which files changed and choosing the smallest relevant test set. Changed paths are matched against a human-authored test-map.json. Each rule names a risk class and one or more logical test families. The selector expands those families from the live test inventory, estimates their cost, and either runs them or refuses.

Changed pathsstaged files or a Git range
→
Impact maprisk + logical familiesrecent correlated failures may promote up to three suites
→
Costed decisionunder budget: run targeted tests
valid receipts: reuse evidenceover budget or milestone: refuse
Figure 3. The selector produces a decision with reasons. A refusal is an intended result, not an exception to route around.

The map is intentionally boring. A file in the Rust core maps to Cargo suites and, where relevant, the native boundary. A Studio component maps to its unit or component test, plus type checking and linting. A shared wire-format change maps to Swift, TypeScript and Rust parity. A changed test file usually maps to that test alone. New paths are not silently treated as low risk: the map-honesty check rejects any tracked file it cannot classify.

The one adaptive input is equally modest. Recent failures are associated with their top-level directories. A failure increases a suite's score, a later success reduces it, and old signals decay. At most three correlated suites are promoted. This is a corrective memory, not an ML oracle.

Risk changes the depth, not the rhetoric

The selector assigns the highest risk touched by the diff. Risk does not directly mean “run everything.” It describes blast radius and widens the relevant family.

R0Docs, tooling, testsProve the changed thing and the infrastructure around it.
R1Product and UIRun the owning module's behavior, static checks and component coverage.
R2Engine logicRun correlated kernels and the applicable shared fixture corpus.
R3Data boundariesSync, persistence, deletion, schemas, FFI and integrity get the broadest correlated families.
Figure 4. The classes are a naive prior over blast radius. They are explicit, reviewable and versioned with the code.

This makes the selector explainable in a way a probabilistic system would not be. If an iOS sync edit selects 47 discovered sync suites, the output shows the path and rule that caused it. If a native-core edit requires rebuilding the XCFramework before an iOS boundary test, that dependency is visible in the map. If the selection is wrong, the mapping can be reviewed and corrected like any other source file.

The trade-off is that the map can only express couplings we know about. Completeness checks catch unmapped files and unknown suites, but they cannot discover a conceptual dependency nobody recorded. Code review, shared parity fixtures and milestone runs still matter. Selection reduces the price of the ordinary loop; it does not abolish systems thinking.

Receipts turn completed work into reusable evidence

Successful runs produce test receipts. This is basically a caching layer that says: we already did this before, and nothing relevant changed. The details matter because a filename that says “passed” is easy to trust and easy to misuse.

A receipt binds the suite and verdict to a clean Git state, records its duration and summary, and points to an artifact with a verified digest. Dirty, stale, incomplete or mismatched evidence does not satisfy the gate. A changed source fingerprint invalidates build reuse. A missing result bundle invalidates receipts that require it. A failed receipt is history, not permission to skip.

Reusable receipt
  • same clean HEAD
  • same suite or gate leg
  • passing verdict
  • artifact digest matches
≠
“It passed earlier”
  • source has moved
  • working tree is dirty
  • artifact is absent
  • run was incomplete
Figure 5. Reuse depends on identity, result and evidence. It is not a time-based promise that a nearby run was probably good enough.

In measured runs, a focused Swift engine test went from 0.55 seconds of execution to a 0.04-second receipt replay. A branch-selected gate went from 41.34 seconds to 0.26 seconds when repeated against the same source state. The larger win is behavioral: agents and humans stop paying for work merely because a new command happened to ask for it.

Ten minutes is a boundary, not a performance claim

The default budget is 600 seconds. Estimates prefer recent passing receipts and fall back to conservative costs. A cold Xcode path receives a 900-second estimate floor, which means it cannot quietly squeeze under a ten-minute budget on optimism. If the estimate is too high, the selector prints the breakdown and refuses unless an explicit, recorded override is provided.

That distinction is important. The tool does not promise every relevant check will finish in ten minutes. It promises not to start an unbounded amount of work while pretending it is the ordinary feedback loop.

Default evidence · every change

  • mapped unit, component and integration families
  • static checks and cross-engine parity where coupled
  • warm build reuse and valid receipts
  • hard budget with an explained refusal

Milestone evidence · explicit

  • full iOS suite
  • all E2E scenarios, repeated
  • scaled synchronization soaks
  • owner-invoked and resumable
Figure 6. Expensive evidence was not deleted. It was moved behind a boundary that makes its cost and purpose explicit.

The full iOS suite, repeated full E2E, and scaled soaks live behind a separate milestone command that requires an owner invocation. Ordinary tools cannot reach those legs by accident. Changed end-to-end scenarios can still run individually when the change warrants it, and a new or modified scenario can be repeated to demonstrate determinism without paying for every unrelated journey.

What changed in practice

The useful change was not that Korivo now tests less. It is that testing has a shape. A local edit receives local evidence. A wire-format edit crosses all three implementations. A synchronization change expands to the real synchronization family. An unknown file stops the gate. A change whose honest blast radius exceeds the default budget is named as milestone work instead of being allowed to consume an hour in the background.

This does introduce risk. A dependency missing from the map may be caught later by CI, an adversarial review, or a milestone run instead of during the first local loop. That is the trade: dramatically cheaper incremental feedback in exchange for accepting that broad regressions can sometimes move to a later gate. The system tries to make that trade visible rather than claiming it away.

I think the result is interesting precisely because it is not sophisticated. There is no service to operate and no model whose reasoning we have to infer. The repository contains the map, the risk classes, the suite registry, the budget, the receipt rules, and tests for the selector itself. When the system makes a bad decision, there is somewhere concrete to fix it.

The memorable lesson was not “7,400 tests are too many.” Most of them were earning their place. The lesson was that evidence has a cost, a scope, and an identity. Once those became explicit, the ordinary loop stopped behaving like a release gate—and the release gate became more meaningful because we knew when we were actually asking for it.