Forecasting backtest
MAE 12.48% (median 9.90%)
Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.
Chaque métrique publiée est rattachée à un protocole reproductible, une spécification matérielle et une marge d'erreur empirique. Aucun chiffre marketing.
MAE 12.48% (median 9.90%)
Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.
71.4% ground-truth accuracy, 0.572 average groundedness
Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.
Confirmed persona-based RBAC actually filters retrieval, not just display
Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.
95.0% (57/60) zero-shot
See protocol.
92.5-100%
See protocol.
77.0% CORD, 100% French/FCFA sample
See protocol.
$0.0007-0.0021/doc (Route B) vs $0.0048-0.0122/doc (Route A)
See protocol.
95.0% (7,488/7,878 records)
Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.
550/550 documents processed successfully (100%) at ~1.1 docs/second
See protocol.
Consensus accuracy 0.785 [0.725, 0.840], F1 0.786 [0.717, 0.843], ROC-AUC 0.870 [0.818, 0.915]
Headline finding, stated plainly: consensus did NOT beat the strongest single judge (Groq gpt-oss-120b at 0.830). A real, disclosed limitation, not a suppressed one.
Stdev 0.272 on wrong predictions vs 0.069 on correct ones
Headline finding, stated plainly: consensus did NOT beat the strongest single judge (Groq gpt-oss-120b at 0.830). A real, disclosed limitation, not a suppressed one.
14/14 enforced correctly, reproducible deterministically with zero network or LLM calls
See protocol.
Average execution 1.8s, P95 3.2s, 20/20 success across all four protocol stages
See protocol.
8.3% (keyword only) → 64.6% (+embedding) → 91.7% (full cascade), 0.793 macro-F1
Throughput ceiling published honestly: a 1,000-request instantaneous burst against a single free-tier instance peaked at 22 req/s with a 100% error rate under that load shape — reported as evidence unthrottled bursts need queuing/backpressure/horizontal scaling, not as a production capacity claim.
100% correct accept/reject (90/90 valid processed, 10/10 invalid rejected)
Throughput ceiling published honestly: a 1,000-request instantaneous burst against a single free-tier instance peaked at 22 req/s with a 100% error rate under that load shape — reported as evidence unthrottled bursts need queuing/backpressure/horizontal scaling, not as a production capacity claim.
2.2% / 0.8%
An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only.
P=0.502, R=0.518, F1=0.506
An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only.
Comparaison des valeurs publiées avec leur contexte d’échantillonnage.