Coexi benchmark results
4.45 ms of pipeline overhead
0.5% structural overhead · 100% recall on gate refusals · 0.0% false positives
Intel Ultra 9 285K · RTX 5090 · 128GB RAM · Windows 11 Pro
Structural overhead
0.5%of full turn E2EFull turn E2E (haiku-4-5)
1065 ms95% CI: ±38 ms · p99: 4310 msGate performance
100% / 0.0%recall / false positive rateLatency profile (mean)
Breakdown of a single turn | Latency (ms) - log scale
Cold start (init, import, model load)
2032.5 ms
Full turn E2E (haiku-4-5)
1065 ms
Pipeline overhead (everything except LLM)
4.45 ms
LOM warm (logic-over-model layer)
0.0031 ms
0.0010.010.11101001,00010,000
Identity (ChromaDB)
Layer 2 mean / p9913.85 ms / 40.6 ms
Layer 3 mean / p9913.69 ms / 47.7 ms
Combined mean13.71 ms
Per-turn identity cost13.71 ms
Cache (p99)
Read0.006 ms
Write0.490 ms
Token throughput
38.3 tok/s
p5: 8.5 tok/sp95: 68.7 tok/s
Prompt cache
98.9%
cache read ratio1000 inference runs
Memory
Idle RSS83.0 MB
Heap alloc / session0.05 KB
Testing
Full turn calls1000 calls
Cold start p992262.6 ms
LOM p990.0046 ms
Samples1,500+
Deterministic accuracy100%
Cold path excluded for warm measurements · 1000 inference runs
Home