SMTRY Labs · research

The road to
sovereign memory.

Labs is where we test claims before they go on a product page. The question under all of it: can memory run entirely inside a network, on a small model, and still be trusted?

August 2026 holdout

In August 2026 we ran a holdout: an 8-billion-parameter open-weight model, quantized to 2 bits, against a cloud arm. Evidence coverage was 0.94 on both arms. The local arm answered 17 of 30 questions correctly; the cloud arm answered 22. It took about 70 hours where the cloud took 6 minutes. The local model can do the work, slowly. Screening before trust runs in the product today, on a cloud model. In a local test, an 8-billion-parameter model returned the same verdicts as the cloud screen on two probes, in 2 to 3 seconds; we have not yet measured agreement at scale.

August 2026 holdout · local 8B 2-bit vs cloud pipeline

Evidence coverage 0.94 local 0.94 cloud Correct answers, of 30 17 local 22 cloud Wall-clock, full corpus ~70 hours, one consumer machine 6 min, cloud batch
One consumer machine, an 8-billion-parameter open-weight model quantized to two bits, distilled the 30 holdout pools and matched the cloud pipeline's retrieval quality exactly. The remaining answer gap is concentrated in temporal reasoning. The trade today is time.

Craft Alpha

Benchmarks

The alpha we run, not sell.

Craft is an alpha we run for our own writing desk rather than sell: a collaborator that learns a writer's instincts and pushes back on the places where habit is standing in for a choice. It edits, pushes back, and keeps a record of what it changed and why. It is not available outside the firm.

Benchmarks

The measured results and their methods are published on the product's benchmarks page and kept current; every measured figure on this site is published there.

Read the benchmarks →