Differential testing. Both runtimes, the same programs, the same inputs —
and every read, every write and every arithmetic result hashed and compared. Not a smoke
test. Not a sample. The whole file, record for record, and the pennies digit for digit under
the same rounding model.
This is exactly how the engine was built in the first place. We did not read a manual — we
measured what the runtime you license today actually does, down to its on-disk
record-locking protocol. That is why both engines can hold the same files open at once
without corrupting each other's work, and it is what turns a parallel run from a slide into
a thing you can watch.
We pointed it at ourselves first. When we rewrote our engine, we ran the new one and the
one it replaces over all seven benchmark programs, twice,
independently, and byte-compared the output: identical, both runs, zero
mismatches. Behind that sit more than 2,200 automated tests, and nearly half of our
source is test code. We kept the harness after it went green, and it is now the migration tool.
So the mirror isn't a demo environment. It is a continuous, machine-checked argument about
whether the two systems are the same system — and you can watch it run for as long as you
like. Most people want at least one clean month-end close diffed before they will talk about
a date, which is the correct instinct.
Byte-exact or it doesn't ship. There is no second definition of "working" here.