NathanRoll commited on
Commit
b3421ca
·
verified ·
1 Parent(s): 76f5e73

Document full 25-language Core ML qualification and current GPT comparison

Browse files
coreml/QUALIFICATION-20260924.md ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Orukeet Core ML qualification — September 24, 2026
2
+
3
+ **Status: full paired accuracy evaluation completed; experimental release with known recognition limitations. Physical iPhone qualification remains pending.** The recognizer stays Orukeet. This report compares the existing INT8 encoder candidate with the existing LUT6 Orukeet baseline and supplies a separate small, current GPT-4o Transcribe comparison. No weights were changed.
4
+
5
+ **Known recognition-quality limitation: Slovenian script errors remain.** On all 834 Slovenian recordings, 23 INT8 outputs and 20 LUT6 outputs contain Cyrillic; 16 and 14 are majority-Cyrillic. Installer/runtime fixes did not repair this recognition issue. Collecting the complete corpus is not complete accuracy acceptance.
6
+
7
+ ## Full 25-language FLEURS test coverage
8
+
9
+ Each Core ML artifact attempted **20,146 recordings (64.71 hours)**, including all 3,520 recordings longer than 15 seconds and all 64 longer than 30 seconds. There was no duration, reference or quality filtering. All 20,146 waveform hashes are distinct. Source metadata, original WAV hashes, decoded sample counts and exact row coverage were verified for both artifacts.
10
+
11
+ The **primary attempted-set** scores are **13.6403% WER for INT8** (56,477/414,045 word errors) and **14.2656% for LUT6** (59,066/414,045). Character error rates are 5.4249% and 5.6609%. INT8 succeeded on 20,146/20,146; LUT6 succeeded on 20,144/20,146. The primary LUT6 score includes the two administrative-interruption failures documented below. Failed observations are empty hypotheses, with reference denominators preserved.
12
+
13
+ | Language | Clips | Reference words | INT8 WER | LUT6 WER | INT8 − LUT6 | INT8 CER | LUT6 CER |
14
+ |---|---:|---:|---:|---:|---:|---:|---:|
15
+ | Bulgarian | 658 | 13,887 | 13.091% | 13.790% | -0.698 pp | 4.918% | 5.229% |
16
+ | Czech | 723 | 13,540 | 11.898% | 12.467% | -0.569 pp | 4.809% | 4.953% |
17
+ | Danish | 930 | 19,981 | 18.528% | 19.333% | -0.806 pp | 7.393% | 7.692% |
18
+ | German | 862 | 18,323 | 7.379% | 7.799% | -0.420 pp | 3.728% | 3.927% |
19
+ | Greek | 650 | 14,864 | 35.159% | 36.652% | -1.494 pp | 11.348% | 12.293% |
20
+ | English | 647 | 14,268 | 6.427% | 6.567% | -0.140 pp | 3.622% | 3.668% |
21
+ | Spanish | 908 | 23,118 | 5.100% | 5.173% | -0.074 pp | 3.014% | 2.982% |
22
+ | Estonian | 893 | 14,656 | 16.717% | 17.767% | -1.051 pp | 4.840% | 5.294% |
23
+ | Finnish | 918 | 14,525 | 13.377% | 14.017% | -0.640 pp | 4.248% | 4.454% |
24
+ | French | 676 | 17,810 | 7.378% | 7.777% | -0.399 pp | 3.660% | 3.834% |
25
+ | Croatian | 914 | 17,368 | 14.037% | 14.705% | -0.668 pp | 5.934% | 6.352% |
26
+ | Hungarian | 905 | 16,705 | 15.660% | 16.702% | -1.042 pp | 5.189% | 5.593% |
27
+ | Italian | 865 | 20,779 | 4.659% | 4.904% | -0.245 pp | 3.013% | 3.144% |
28
+ | Lithuanian | 986 | 17,013 | 20.537% | 21.736% | -1.199 pp | 6.996% | 7.300% |
29
+ | Latvian | 851 | 15,271 | 22.631% | 24.052% | -1.421 pp | 7.094% | 7.289% |
30
+ | Maltese | 926 | 21,750 | 24.611% | 25.131% | -0.520 pp | 10.388% | 10.832% |
31
+ | Dutch | 364 | 8,251 | 8.605% | 8.847% | -0.242 pp | 4.046% | 4.165% |
32
+ | Polish | 758 | 14,253 | 8.784% | 9.177% | -0.393 pp | 4.019% | 4.145% |
33
+ | Portuguese | 919 | 21,329 | 6.447% | 6.601% | -0.155 pp | 3.783% | 3.855% |
34
+ | Romanian | 883 | 20,776 | 12.312% | 13.294% | -0.982 pp | 5.032% | 5.324% |
35
+ | Russian | 775 | 14,927 | 8.602% | 8.870% | -0.268 pp | 3.991% | 4.057% |
36
+ | Slovak | 792 | 14,942 | 10.795% | 11.264% | -0.468 pp | 4.336% | 4.534% |
37
+ | Slovenian | 834 | 16,264 | 27.429% | 28.185% | -0.756 pp | 11.040% | 11.018% |
38
+ | Swedish | 759 | 15,086 | 15.034% | 15.869% | -0.835 pp | 5.654% | 5.956% |
39
+ | Ukrainian | 750 | 14,359 | 7.793% | 8.162% | -0.369 pp | 3.048% | 3.129% |
40
+
41
+ Lower is better. WER is ordinary token Levenshtein after the same normalization, with a fixed reference-word denominator for both models. CER includes normalized spaces. Every primary word and character distance was independently checked with Kaldialign. Pair-dependent compound WER is only a secondary diagnostic and is not used for this table.
42
+
43
+ | Duration subset | Clips per artifact | INT8 WER | LUT6 WER |
44
+ |---|---:|---:|---:|
45
+ | Longer than 15 seconds | 3,520 | 14.407% | 15.001% |
46
+ | Longer than 30 seconds | 64 | 13.433% | 14.346% |
47
+
48
+ The recordings span 2.70–52.92 seconds. The harness uses the production engine’s normal long-recording path without manually splitting, padding or truncating input. This is not a minutes-long continuous-dictation accuracy test.
49
+
50
+ ## Administrative interruption accounting
51
+
52
+ A coordinated CPU handoff suspended the active benchmark processes. Two LUT6 predictions timed out at exactly the in-flight Hungarian and Italian rows recorded before the pause outcomes were inspected; each request included approximately 581 seconds of suspension. The corresponding INT8 in-flight requests succeeded, and subsequent inference continued. These are separately identified administrative-pause failures, not evidence of two intrinsic uninterrupted-runtime failures. Their empty hypotheses remain in the primary table.
53
+
54
+ After the main run, all four pre-recorded in-flight fixture IDs were replayed under **both** artifacts: eight additional, separately retained observations. The secondary view below substitutes only those exact four rows per model. Every other observation and every denominator is unchanged; the original failure records are retained. **This confirmatory substituted score is not the primary result.**
55
+
56
+ | Secondary pause-control view | Word errors / reference words | WER | CER | Successful / attempted |
57
+ |---|---:|---:|---:|---:|
58
+ | INT8 | 56,477 / 414,045 | 13.6403% | 5.4249% | 20,146 / 20,146 |
59
+ | LUT6 | 59,020 / 414,045 | 14.2545% | 5.6501% | 20,146 / 20,146 |
60
+
61
+ The four fixture IDs are `hu_hu/1371634410209825983.wav`, `hu_hu/1591928531702704285.wav`, `it_it/12715874722470470383.wav` and `it_it/12013566675598043608.wav`. None belongs to the 400-clip GPT comparison below.
62
+
63
+ ## Known Slovenian script failures
64
+
65
+ Across all 834 Slovenian test recordings, any Cyrillic appears in **23 INT8 outputs and 20 LUT6 outputs**. Majority-Cyrillic counts are 16 and 14. These Unicode flags identify transcripts for review; no flagged result was removed, transliterated or corrected before scoring.
66
+
67
+ The two previously exposed recordings were reproduced with the same engine using all three existing precision profiles. “Cyrillic” below means any Cyrillic code point appears in the output; no transcript text is reproduced here.
68
+
69
+ | Previously exposed Slovenian recording | LUT6 Cyrillic | INT8 Cyrillic | FP16 Cyrillic |
70
+ |---|---|---|---|
71
+ | `10566187729969396496.wav` | Yes | Yes | Yes |
72
+ | `10749956690098540209.wav` | No | Yes | Yes |
73
+
74
+ INT8/LUT6 reproduction transcripts matched the full run. FP16 also reproducing a script failure means it is not established as an INT8-only defect. These observations do not determine the complete root cause, and pooled WER improvement does not resolve a language/script error. No script-based repair or decoding adjustment was applied.
75
+
76
+ ## Fresh 400-clip GPT-4o Transcribe comparison
77
+
78
+ The complete diagnostic panel frozen on September 20 contains **16 clips per language, 400 total, 68.04 audio minutes**, all at most 15 seconds. The September 24 GPT file-API run completed 400/400 requests without retries. The current Core ML observations are joined by exact WAV hash, source sample count and reference. Primary WER is **7.451% GPT-4o** (570/7,650), **14.013% Orukeet INT8** (1,072/7,650) and **14.536% Orukeet LUT6** (1,112/7,650).
79
+
80
+ | Language | Clips | Reference words | GPT-4o WER | Orukeet INT8 WER | Orukeet LUT6 WER | INT8 − GPT |
81
+ |---|---:|---:|---:|---:|---:|---:|
82
+ | Bulgarian | 16 | 310 | 10.00% | 15.48% | 16.45% | +5.48 pp |
83
+ | Czech | 16 | 238 | 3.78% | 12.61% | 12.61% | +8.82 pp |
84
+ | Danish | 16 | 294 | 8.84% | 23.13% | 23.81% | +14.29 pp |
85
+ | German | 16 | 398 | 2.26% | 5.53% | 5.78% | +3.27 pp |
86
+ | Greek | 16 | 343 | 7.00% | 33.53% | 36.15% | +26.53 pp |
87
+ | English | 16 | 343 | 4.08% | 4.96% | 4.66% | +0.87 pp |
88
+ | Spanish | 16 | 372 | 2.96% | 4.03% | 4.03% | +1.08 pp |
89
+ | Estonian | 16 | 211 | 9.95% | 18.96% | 21.80% | +9.00 pp |
90
+ | Finnish | 16 | 236 | 2.54% | 12.29% | 12.29% | +9.75 pp |
91
+ | French | 16 | 383 | 5.22% | 8.36% | 8.88% | +3.13 pp |
92
+ | Croatian | 16 | 313 | 16.29% | 8.63% | 9.58% | -7.67 pp |
93
+ | Hungarian | 16 | 266 | 9.77% | 19.55% | 19.92% | +9.77 pp |
94
+ | Italian | 16 | 351 | 3.42% | 5.41% | 5.41% | +1.99 pp |
95
+ | Lithuanian | 16 | 284 | 11.27% | 22.54% | 24.65% | +11.27 pp |
96
+ | Latvian | 16 | 254 | 5.51% | 18.50% | 22.05% | +12.99 pp |
97
+ | Maltese | 16 | 296 | 23.99% | 20.27% | 19.26% | -3.72 pp |
98
+ | Dutch | 16 | 348 | 6.03% | 8.05% | 7.76% | +2.01 pp |
99
+ | Polish | 16 | 279 | 5.02% | 11.83% | 12.54% | +6.81 pp |
100
+ | Portuguese | 16 | 333 | 6.61% | 7.81% | 7.21% | +1.20 pp |
101
+ | Romanian | 16 | 356 | 2.81% | 13.76% | 15.17% | +10.96 pp |
102
+ | Russian | 16 | 293 | 5.80% | 7.85% | 7.51% | +2.05 pp |
103
+ | Slovak | 16 | 303 | 5.61% | 14.52% | 13.86% | +8.91 pp |
104
+ | Slovenian | 16 | 270 | 17.41% | 42.22% | 41.85% | +24.81 pp |
105
+ | Swedish | 16 | 297 | 12.79% | 18.86% | 19.53% | +6.06 pp |
106
+ | Ukrainian | 16 | 279 | 2.51% | 5.02% | 5.02% | +2.51 pp |
107
+
108
+ This is a small descriptive comparison, not a reliable language-winner ranking. The Core ML comparator is the mobile batch runtime, not the current cloud-native Orukeet service. It is separate from any GPT-OSS text-cleanup experiment. No automatic provider routing or recognizer replacement follows from this table.
109
+
110
+ The API requested `gpt-4o-transcribe`, JSON response and temperature zero, without a prompt, language parameter, reference text or prior transcript. Multipart filenames retained the existing language-prefixed diagnostic basenames; all 400 responses reported zero text-input tokens. The alias is floating and no immutable provider snapshot was returned. Modeled usage cost was **$0.2582**, calculated from returned usage and [published pricing](https://developers.openai.com/api/docs/pricing), not an invoice. The API calls were paced; no latency or speed comparison is supported. [Request schema](https://developers.openai.com/api/reference/cli/resources/audio/subresources/transcriptions/methods/create).
111
+
112
+ ## Identity, reproducibility and release scope
113
+
114
+ - Dataset: [Google FLEURS](https://huggingface.co/datasets/google/fleurs/tree/70bb2e84b976b7e960aa89f1c648e09c59f894dd), revision `70bb2e84b976b7e960aa89f1c648e09c59f894dd`, original test WAVs and references, CC-BY-4.0.
115
+ - Runtime source: Orukeet Engine/Audio commit `db30e0b1c27ea0c88ae789473fff6704fc3b4445`; FluidAudio commit `ddc95f4d03d5be12bf84eeb6e4bde5b724356d88`. The current SDK runtime/audio files were verified byte-identical to those qualified sources; installer changes are a separate check.
116
+ - Execution: Apple M5 Max, 128 GiB unified memory, macOS 26.4.1, Swift 6.2.4. Encoder/decoder/joint use CPU + Neural Engine; preprocessor uses CPU. Four independent accuracy workers, each with production batch concurrency four. Concurrent/suspended run times are not performance claims.
117
+ - INT8 portable archive SHA-256: `24df9ff76f00f86f9ae1fd601cbbcab1d1eac98c7e8107de67444a7858d88b8b`. All 18 expected compiled files per artifact were freshly hashed; all four INT8 component weights, compiled representations and vocabulary match the existing portable-source correspondence receipt.
118
+ - Normalizer source revision: `852c3e355f20a111a2ee39f76677bc0ba65147bf`. Both models use identical normalized references and fixed denominators. Detailed metadata, hashes, request journals, attempt records and independent arithmetic receipts are retained for audit.
119
+ - Full FLEURS closes the missing all-duration conversion evaluation. It has model-development exposure and is not an unseen generalization benchmark. No acceptable-regression threshold was declared beforehand, so corpus completion must not be relabeled universal accuracy acceptance.
120
+ - This does not establish superiority to other Parakeet v2/v3 Core ML conversions, benchmark continuous streaming, validate microphone capture on a physical iPhone, or qualify cold installation, memory/thermal behavior, backgrounding, cancellation and repeated sessions on the supported device floor.
121
+ - Keep the current release experimental while the known script errors and physical-device checks are resolved. Model identity remains Orukeet; mobile input remains 16 kHz mono with record-then-transcribe batch inference.
122
+
123
+ Evaluation receipts completed 2026-09-24T19:56:25.887944+00:00; the separate interruption confirmation completed 2026-09-24T19:56:31.680648+00:00.
coreml/README.md CHANGED
@@ -21,11 +21,18 @@ The installed cache is approximately **1.19 GB**, including the retained ZIP; fi
21
  peak is approximately **1.82 GB** before an app-specific free-space margin. Compilation size
22
  varies by device and OS. Legacy installs without a retained source require one fresh download.
23
 
24
- This candidate is experimental and opt-in. The original 400-clip diagnostic panel does not
25
- establish 25-language production acceptance; it contains language/script regressions, including
26
- Slovenian Cyrillic output. Full-corpus qualification is in progress, and physical iPhone
27
- performance/thermal/microphone acceptance remains required. The archive itself is immutable;
28
- qualification updates are separate metadata, not modifications of its weights or decoder.
 
 
 
 
 
 
 
29
 
30
  [Integration and device acceptance guide](https://github.com/Oruk-AI/orukeet/blob/v0.1.2-coreml.1/integrations/openwhispr/ios/README.md)
31
 
@@ -119,8 +126,8 @@ The baseline uses 6-bit LUT/FP16 weights; the greedy profile removes unused
119
  top-64 joint calculations. It is unsuitable for decoding paths that require
120
  those outputs. The separate iOS experimental precision profile is described above.
121
 
122
- This is a preview. Other Macs, older OS versions, power use, and the full
123
- 25-language Core ML evaluation remain unqualified. FluidAudio's default
124
  sliding-window TDT path buffers roughly 13 seconds before first text. This is
125
  not the separate Parakeet EOU 120M live-typing model.
126
 
 
21
  peak is approximately **1.82 GB** before an app-specific free-space margin. Compilation size
22
  varies by device and OS. Legacy installs without a retained source require one fresh download.
23
 
24
+ This candidate remains experimental and opt-in. The full paired 25-language evaluation is
25
+ complete: 20,146 recordings per artifact, including long inputs. INT8 WER is **13.6403%**
26
+ versus **14.2656%** for LUT6 on 414,045 reference words. The primary LUT6 score preserves
27
+ two documented administrative-pause timeouts; its separate replay-control WER is 14.2545%.
28
+ Known Slovenian script errors remain (Cyrillic in 23/834 INT8 and 20/834 LUT6 outputs).
29
+ The fresh, separate 400-clip comparison reports GPT-4o Transcribe 7.451% WER versus
30
+ Orukeet INT8 14.013%; it is a small descriptive mobile-runtime panel, not a cloud benchmark.
31
+
32
+ [Full qualification report and per-language tables](QUALIFICATION-20260924.md)
33
+
34
+ Physical iPhone performance, thermal and microphone acceptance remains required. The archive
35
+ itself is immutable; qualification updates are separate metadata, with no weight or decoder changes.
36
 
37
  [Integration and device acceptance guide](https://github.com/Oruk-AI/orukeet/blob/v0.1.2-coreml.1/integrations/openwhispr/ios/README.md)
38
 
 
126
  top-64 joint calculations. It is unsuitable for decoding paths that require
127
  those outputs. The separate iOS experimental precision profile is described above.
128
 
129
+ These historical macOS measurements do not qualify other devices, older OS versions or
130
+ power use. The current full-corpus INT8/LUT6 evaluation is documented separately above. FluidAudio's default
131
  sliding-window TDT path buffers roughly 13 seconds before first text. This is
132
  not the separate Parakeet EOU 120M live-typing model.
133
 
coreml/manifest.json CHANGED
@@ -37,8 +37,33 @@
37
  "fluid_audio_commit": "ddc95f4d03d5be12bf84eeb6e4bde5b724356d88",
38
  "source_pr": "https://github.com/Oruk-AI/orukeet/pull/10",
39
  "qualification": {
40
- "full_fleurs": "in_progress",
41
- "physical_iphone": "pending"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
  },
43
  "source_code_revision": "6c37c587fabcef8b0e932584fbf22a79fcb9d89e",
44
  "sdk_release_tag": "v0.1.2-coreml.1",
 
37
  "fluid_audio_commit": "ddc95f4d03d5be12bf84eeb6e4bde5b724356d88",
38
  "source_pr": "https://github.com/Oruk-AI/orukeet/pull/10",
39
  "qualification": {
40
+ "full_fleurs": "completed_with_known_recognition_limitations",
41
+ "physical_iphone": "pending",
42
+ "report": "QUALIFICATION-20260924.md",
43
+ "evaluated_at": "2026-09-24",
44
+ "languages": 25,
45
+ "recordings_per_artifact": 20146,
46
+ "reference_words": 414045,
47
+ "int8_word_errors": 56477,
48
+ "lut6_word_errors_primary": 59066,
49
+ "int8_wer_percent": 13.6403,
50
+ "lut6_wer_percent_primary": 14.2656,
51
+ "lut6_wer_percent_pause_control_secondary": 14.2545,
52
+ "int8_inference_failures": 0,
53
+ "lut6_administrative_pause_failures": 2,
54
+ "slovenian_cyrillic_outputs": {
55
+ "int8": 23,
56
+ "lut6": 20,
57
+ "recordings_per_artifact": 834
58
+ },
59
+ "gpt_comparison": {
60
+ "clips": 400,
61
+ "reference_words": 7650,
62
+ "gpt4o_wer_percent": 7.451,
63
+ "int8_wer_percent": 14.0131,
64
+ "scope": "small_descriptive_mobile_runtime_panel"
65
+ },
66
+ "status": "experimental_opt_in"
67
  },
68
  "source_code_revision": "6c37c587fabcef8b0e932584fbf22a79fcb9d89e",
69
  "sdk_release_tag": "v0.1.2-coreml.1",