Yaçine Seybou Siddo
AI Systems Engineer & Independent Consultant
|
Accueil/Benchmarks Empiriques
Télémétrie de Production Vérifiée

Benchmarks

Chaque métrique publiée est rattachée à un protocole reproductible, une spécification matérielle et une marge d'erreur empirique. Aucun chiffre marketing.

Forecasting backtest

MAE 12.48% (median 9.90%)

Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.

Télécharger JSON

Production RAG accuracy

71.4% ground-truth accuracy, 0.572 average groundedness

Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.

Télécharger JSON

A/B test

Confirmed persona-based RBAC actually filters retrieval, not just display

Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.

Télécharger JSON

SROIE public benchmark

95.0% (57/60) zero-shot

See protocol.

Télécharger JSON

Route A accuracy

92.5-100%

See protocol.

Télécharger JSON

Route B accuracy

77.0% CORD, 100% French/FCFA sample

See protocol.

Télécharger JSON

Cost

$0.0007-0.0021/doc (Route B) vs $0.0048-0.0122/doc (Route A)

See protocol.

Télécharger JSON

GraphRAG-lite entity coverage

95.0% (7,488/7,878 records)

Bilingual parity check: French scored 0.917 vs. English 0.431 on the same 50-case run — a real, disclosed gap, not smoothed over.

Télécharger JSON

Throughput/reliability

550/550 documents processed successfully (100%) at ~1.1 docs/second

See protocol.

Télécharger JSON

HaluEval-QA validation

Consensus accuracy 0.785 [0.725, 0.840], F1 0.786 [0.717, 0.843], ROC-AUC 0.870 [0.818, 0.915]

Headline finding, stated plainly: consensus did NOT beat the strongest single judge (Groq gpt-oss-120b at 0.830). A real, disclosed limitation, not a suppressed one.

Télécharger JSON

Panel disagreement signal

Stdev 0.272 on wrong predictions vs 0.069 on correct ones

Headline finding, stated plainly: consensus did NOT beat the strongest single judge (Groq gpt-oss-120b at 0.830). A real, disclosed limitation, not a suppressed one.

Télécharger JSON

Adversarial guardrail cases

14/14 enforced correctly, reproducible deterministically with zero network or LLM calls

See protocol.

Télécharger JSON

MCP tool-invocation benchmark

Average execution 1.8s, P95 3.2s, 20/20 success across all four protocol stages

See protocol.

Télécharger JSON

Classification cascade accuracy

8.3% (keyword only) → 64.6% (+embedding) → 91.7% (full cascade), 0.793 macro-F1

Throughput ceiling published honestly: a 1,000-request instantaneous burst against a single free-tier instance peaked at 22 req/s with a 100% error rate under that load shape — reported as evidence unthrottled bursts need queuing/backpressure/horizontal scaling, not as a production capacity claim.

Télécharger JSON

Webhook HMAC verification

100% correct accept/reject (90/90 valid processed, 10/10 invalid rejected)

Throughput ceiling published honestly: a 1,000-request instantaneous burst against a single free-tier instance peaked at 22 req/s with a 100% error rate under that load shape — reported as evidence unthrottled bursts need queuing/backpressure/horizontal scaling, not as a production capacity claim.

Télécharger JSON

WER/CER

2.2% / 0.8%

An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only.

Télécharger JSON

Action-item extraction

P=0.502, R=0.518, F1=0.506

An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only.

Télécharger JSON

Visualisations benchmark

Comparaison des valeurs publiées avec leur contexte d’échantillonnage.

0550
Forecasting backtestMAE 12.48% (median 9.90%)
Production RAG accuracy71.4% ground-truth accuracy, 0.572 average groundedness
SROIE public benchmark95.0% (57/60) zero-shot
Route A accuracy92.5-100%
Route B accuracy77.0% CORD, 100% French/FCFA sample
Cost$0.0007-0.0021/doc (Route B) vs $0.0048-0.0122/doc (Route A)
GraphRAG-lite entity coverage95.0% (7,488/7,878 records)
Throughput/reliability550/550 documents processed successfully (100%) at ~1.1 docs/second
HaluEval-QA validationConsensus accuracy 0.785 [0.725, 0.840], F1 0.786 [0.717, 0.843], ROC-AUC 0.870 [0.818, 0.915]
Panel disagreement signalStdev 0.272 on wrong predictions vs 0.069 on correct ones
Adversarial guardrail cases14/14 enforced correctly, reproducible deterministically with zero network or LLM calls
MCP tool-invocation benchmarkAverage execution 1.8s, P95 3.2s, 20/20 success across all four protocol stages
Classification cascade accuracy8.3% (keyword only) → 64.6% (+embedding) → 91.7% (full cascade), 0.793 macro-F1
Webhook HMAC verification100% correct accept/reject (90/90 valid processed, 10/10 invalid rejected)
WER/CER2.2% / 0.8%
Action-item extractionP=0.502, R=0.518, F1=0.506