Automatic Speech Recognition
NeMo
ONNX
Safetensors
GGUF
parakeet_tdt
parakeet
tdt
sherpa-onnx
multilingual
speech-recognition
gabor
fastconformer
Instructions to use oruk/orukeet with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use oruk/orukeet with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("oruk/orukeet") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
File size: 8,400 Bytes
76f5e73 b3421ca fe3bbba 7df0b47 b3421ca 76f5e73 b59c13a 76f5e73 43142dd 76f5e73 43142dd b3421ca 43142dd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 | # Orukeet r3 Core ML previews
Portable Core ML bundles for Apple Silicon. The original macOS previews below use **FluidAudio 0.15.5**.
The iOS 17+ opt-in candidate uses the pinned Oruk FluidAudio fork through OrukeetCoreML.
## iOS / OpenWhispr candidate
The experimental INT8 encoder archive is published for opt-in integration with
[OrukeetCoreML v0.1.2-coreml.1](https://github.com/Oruk-AI/orukeet/releases/tag/v0.1.2-coreml.1). Audio is 16 kHz mono;
record first, then run local batch inference. Keep the engine loaded between recordings.
- [Download the INT8 archive](https://huggingface.co/oruk/orukeet/resolve/419d7f79e290127e202a0f610509868d314743eb/coreml/orukeet-r3-coreml-int8sym-encoder-only-experimental-20260920-deflated.zip?download=true)
- Immutable model revision: `419d7f79e290127e202a0f610509868d314743eb`
- Download bytes: **554,985,744**
- SHA-256: `24df9ff76f00f86f9ae1fd601cbbcab1d1eac98c7e8107de67444a7858d88b8b`
- Archive root: `orukeet-r3-coreml-int8sym-encoder-only-experimental-20260920/`
Use the SDK model store to authenticate, safely extract and compile packages on-device.
It retains the authenticated source ZIP for offline recompilation after an OS update.
The installed cache is approximately **1.19 GB**, including the retained ZIP; first-install
peak is approximately **1.82 GB** before an app-specific free-space margin. Compilation size
varies by device and OS. Legacy installs without a retained source require one fresh download.
This candidate remains experimental and opt-in. The full paired 25-language evaluation is
complete: 20,146 recordings per artifact, including long inputs. INT8 WER is **13.6403%**
versus **14.2656%** for LUT6 on 414,045 reference words. The primary LUT6 score preserves
two documented administrative-pause timeouts; its separate replay-control WER is 14.2545%.
Known Slovenian script errors remain (Cyrillic in 23/834 INT8 and 20/834 LUT6 outputs).
The fresh, separate 400-clip comparison reports GPT-4o Transcribe 7.451% WER versus
Orukeet INT8 14.013%; it is a small descriptive mobile-runtime panel, not a cloud benchmark.
[Full qualification report and per-language tables](QUALIFICATION-20260924.md)
The consumer's native inference sequence was also checked separately: 490/490 observations
matched the SDK, including all 64 recordings longer than 30 seconds. The Orukeet-only
preparation warmup preserved every observed transcript, token timestamp and duration.
[Consumer call-path and warmup report](CONSUMER-PARITY-20260924.md). This bounded Mac
check does not replace physical-iPhone or complete-app lifecycle acceptance.
A separate matched-PCM three-case [native/CoreML diagnostic](NATIVE-SLOVENIAN-DIAGNOSTIC-20260924.md)
reproduces one known Slovenian script failure in the native BF16 service and all CoreML
profiles; a second varies across paths. Neither a complete cause nor an accuracy repair
is established. Original corpus and consumer-parity scores are unchanged.
Physical iPhone performance, thermal and microphone acceptance remains required. The archive
itself is immutable; qualification updates are separate metadata, with no weight or decoder changes.
[Integration and device acceptance guide](https://github.com/Oruk-AI/orukeet/blob/decfb6e4b50a41fbddadf8a9cf12beadf06ae2f1/integrations/openwhispr/ios/README.md)
## Original macOS previews
These are the same archives as the
[GitHub preview release](https://github.com/Oruk-AI/orukeet/releases/tag/coreml-taptalk-preview-20260915),
published here for TapTalk and other applications to download.
| Archive | Profile | Bytes |
|:--|:--|--:|
| [orukeet-r3-coreml-greedy.zip](https://huggingface.co/oruk/orukeet/resolve/coreml-taptalk-preview-20260915/coreml/orukeet-r3-coreml-greedy.zip?download=true) | Recommended for ordinary greedy decoding | 466,579,943 |
| [orukeet-r3-coreml-baseline.zip](https://huggingface.co/oruk/orukeet/resolve/coreml-taptalk-preview-20260915/coreml/orukeet-r3-coreml-baseline.zip?download=true) | Retains top-64 outputs for language hints/reranking | 466,579,851 |
[SHA256SUMS.txt](https://huggingface.co/oruk/orukeet/resolve/coreml-taptalk-preview-20260915/coreml/SHA256SUMS.txt)
contains the archive hashes. Each archive includes a `bundle.json` with per-file
SHA-256 hashes, conversion receipts, attribution, and the weight license.
## Download
```sh
python -m pip install 'huggingface-hub>=0.34,<2'
```
```python
import hashlib
from pathlib import Path
from zipfile import ZipFile
from huggingface_hub import hf_hub_download
archive = Path(hf_hub_download(
repo_id="oruk/orukeet",
filename="coreml/orukeet-r3-coreml-greedy.zip",
revision="coreml-taptalk-preview-20260915",
))
expected = "beccdc6f18c4b10527a764f6e3ab12e3e11b969220c0cee175b3bb7eaa94290e"
with archive.open("rb") as stream:
digest = hashlib.sha256()
for block in iter(lambda: stream.read(8 * 1024 * 1024), b""):
digest.update(block)
if digest.hexdigest() != expected:
raise ValueError("Core ML archive checksum mismatch")
with ZipFile(archive) as bundle:
bundle.extractall("models")
print(Path("models/orukeet-r3-coreml-greedy").resolve())
```
For an application downloader, use the archive URL above and verify its hash
before extraction. The archive has one top-level directory,
`orukeet-r3-coreml-greedy/` (or `orukeet-r3-coreml-baseline/`).
## Load with Swift
Each bundle contains `Preprocessor.mlpackage`, `Encoder.mlpackage`,
`Decoder.mlpackage`, `JointDecisionv3.mlpackage`, and `parakeet_vocab.json`.
Compile all four packages on the destination Mac during installation;
machine-specific `.mlmodelc` caches are deliberately excluded.
The `OrukeetCoreML` library is in `export/coreml/benchmark` at
[Oruk-AI/orukeet commit 347f646](https://github.com/Oruk-AI/orukeet/tree/347f646cacda2e001865b6ac40ba5cbc7e90d1c9/export/coreml/benchmark).
Add that directory as a local Swift package dependency, then:
```swift
import OrukeetCoreML
try OrukeetLocalModels.compilePackages(from: downloadedBundle, to: installedCache)
let engine = OrukeetEngine(modelDirectory: installedCache)
try await engine.ensureLoaded()
let result = try await engine.transcribe(samples: mono16kSamples)
print(result.text)
```
`downloadedBundle` is the extracted profile directory; `installedCache` is a
fresh app-owned directory. Audio must be mono 16 kHz PCM. Keep the engine loaded
between recordings and keep the Orukeet cache separate from Parakeet's.
Use the explicit local loader rather than FluidAudio's NVIDIA model downloader.
[Full conversion and integration guide](https://github.com/Oruk-AI/orukeet/blob/347f646cacda2e001865b6ac40ba5cbc7e90d1c9/export/coreml/README.md)
· [Source PR #6](https://github.com/Oruk-AI/orukeet/pull/6)
· [TapTalk](https://github.com/vakharwalad23/tap-talk)
## Measurements and scope
On one M5 Max, greedy decoding reduced paired batch latency by 9.5% relative
to the Core ML baseline. Baseline and greedy returned identical text on 128
FLEURS recordings across eight languages. Core ML WER was 7.64%, compared with
8.55% for Parakeet Core ML and 7.29% for the uncompressed Orukeet source.
[Raw evidence and methodology](https://github.com/Oruk-AI/orukeet/tree/347f646cacda2e001865b6ac40ba5cbc7e90d1c9/evidence/coreml-taptalk-20260915).
The baseline uses 6-bit LUT/FP16 weights; the greedy profile removes unused
top-64 joint calculations. It is unsuitable for decoding paths that require
those outputs. The separate iOS experimental precision profile is described above.
These historical macOS measurements do not qualify other devices, older OS versions or
power use. The current full-corpus INT8/LUT6 evaluation is documented separately above. FluidAudio's default
sliding-window TDT path buffers roughly 13 seconds before first text. This is
not the separate Parakeet EOU 120M live-typing model.
## License and attribution
Orukeet modified weights: CC BY-SA 4.0, retaining NVIDIA Parakeet attribution.
Core ML graphs derive from Fluid Inference's pinned Parakeet conversion.
FluidAudio runtime: Apache-2.0. Conversion and integration code: MIT.
The bundles retain `LICENSE-WEIGHTS`, `NOTICE.md`, and `COREML-NOTICE.txt`.
All weights derive from the r3 checkpoint with SHA-256
`031c8ddab4845aeced904a7cde8e8aa57993b2e344716cf83a545b079c473b56`.
Baseline graph metadata retains the upstream names for graph identity; use
`bundle.json` to identify the weights as Orukeet.
|