{"id":"system-voiceflow-claim-2","systemId":"system-voiceflow","metricLabel":{"EN":"Action-item extraction","FR":"Action-item extraction"},"value":"P=0.502, R=0.518, F1=0.506","protocol":{"EN":"Claude Sonnet 4.6, across 38 of 50 meetings completing a real synthesis→transcription→extraction pipeline end to end; 12 timeouts itemized, not dropped","FR":"Claude Sonnet 4.6, across 38 of 50 meetings completing a real synthesis→transcription→extraction pipeline end to end; 12 timeouts itemized, not dropped"},"source":"https://github.com/Yacine-ai-tech/voiceflow","limitation":{"EN":"An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only.","FR":"An earlier, more optimistic N=20 WER figure (2.9%) was retained and explicitly labeled as such rather than discarded. A prior benchmark script found reporting synthetic, formula-derived WER/CER (not measured) was caught and removed; the corrected report now publishes only what's actually measured. Never framed as beating state-of-the-art — stated as a controlled benchmark result only."},"verified":true,"permalink":"/api/benchmarks/system-voiceflow-claim-2"}