orukeet / coreml /README.md
NathanRoll's picture
Point integration guide at qualified documentation revision
b59c13a verified
|
Raw History Blame Contribute Delete
8.4 kB

Orukeet r3 Core ML previews

Portable Core ML bundles for Apple Silicon. The original macOS previews below use FluidAudio 0.15.5. The iOS 17+ opt-in candidate uses the pinned Oruk FluidAudio fork through OrukeetCoreML.

iOS / OpenWhispr candidate

The experimental INT8 encoder archive is published for opt-in integration with OrukeetCoreML v0.1.2-coreml.1. Audio is 16 kHz mono; record first, then run local batch inference. Keep the engine loaded between recordings.

  • Download the INT8 archive
  • Immutable model revision: 419d7f79e290127e202a0f610509868d314743eb
  • Download bytes: 554,985,744
  • SHA-256: 24df9ff76f00f86f9ae1fd601cbbcab1d1eac98c7e8107de67444a7858d88b8b
  • Archive root: orukeet-r3-coreml-int8sym-encoder-only-experimental-20260920/

Use the SDK model store to authenticate, safely extract and compile packages on-device. It retains the authenticated source ZIP for offline recompilation after an OS update. The installed cache is approximately 1.19 GB, including the retained ZIP; first-install peak is approximately 1.82 GB before an app-specific free-space margin. Compilation size varies by device and OS. Legacy installs without a retained source require one fresh download.

This candidate remains experimental and opt-in. The full paired 25-language evaluation is complete: 20,146 recordings per artifact, including long inputs. INT8 WER is 13.6403% versus 14.2656% for LUT6 on 414,045 reference words. The primary LUT6 score preserves two documented administrative-pause timeouts; its separate replay-control WER is 14.2545%. Known Slovenian script errors remain (Cyrillic in 23/834 INT8 and 20/834 LUT6 outputs). The fresh, separate 400-clip comparison reports GPT-4o Transcribe 7.451% WER versus Orukeet INT8 14.013%; it is a small descriptive mobile-runtime panel, not a cloud benchmark.

Full qualification report and per-language tables

The consumer's native inference sequence was also checked separately: 490/490 observations matched the SDK, including all 64 recordings longer than 30 seconds. The Orukeet-only preparation warmup preserved every observed transcript, token timestamp and duration. Consumer call-path and warmup report. This bounded Mac check does not replace physical-iPhone or complete-app lifecycle acceptance.

A separate matched-PCM three-case native/CoreML diagnostic reproduces one known Slovenian script failure in the native BF16 service and all CoreML profiles; a second varies across paths. Neither a complete cause nor an accuracy repair is established. Original corpus and consumer-parity scores are unchanged.

Physical iPhone performance, thermal and microphone acceptance remains required. The archive itself is immutable; qualification updates are separate metadata, with no weight or decoder changes.

Integration and device acceptance guide

Original macOS previews

These are the same archives as the GitHub preview release, published here for TapTalk and other applications to download.

Archive Profile Bytes
orukeet-r3-coreml-greedy.zip Recommended for ordinary greedy decoding 466,579,943
orukeet-r3-coreml-baseline.zip Retains top-64 outputs for language hints/reranking 466,579,851

SHA256SUMS.txt contains the archive hashes. Each archive includes a bundle.json with per-file SHA-256 hashes, conversion receipts, attribution, and the weight license.

Download

python -m pip install 'huggingface-hub>=0.34,<2'
import hashlib
from pathlib import Path
from zipfile import ZipFile
from huggingface_hub import hf_hub_download

archive = Path(hf_hub_download(
    repo_id="oruk/orukeet",
    filename="coreml/orukeet-r3-coreml-greedy.zip",
    revision="coreml-taptalk-preview-20260915",
))
expected = "beccdc6f18c4b10527a764f6e3ab12e3e11b969220c0cee175b3bb7eaa94290e"
with archive.open("rb") as stream:
    digest = hashlib.sha256()
    for block in iter(lambda: stream.read(8 * 1024 * 1024), b""):
        digest.update(block)
if digest.hexdigest() != expected:
    raise ValueError("Core ML archive checksum mismatch")
with ZipFile(archive) as bundle:
    bundle.extractall("models")
print(Path("models/orukeet-r3-coreml-greedy").resolve())

For an application downloader, use the archive URL above and verify its hash before extraction. The archive has one top-level directory, orukeet-r3-coreml-greedy/ (or orukeet-r3-coreml-baseline/).

Load with Swift

Each bundle contains Preprocessor.mlpackage, Encoder.mlpackage, Decoder.mlpackage, JointDecisionv3.mlpackage, and parakeet_vocab.json. Compile all four packages on the destination Mac during installation; machine-specific .mlmodelc caches are deliberately excluded.

The OrukeetCoreML library is in export/coreml/benchmark at Oruk-AI/orukeet commit 347f646. Add that directory as a local Swift package dependency, then:

import OrukeetCoreML

try OrukeetLocalModels.compilePackages(from: downloadedBundle, to: installedCache)
let engine = OrukeetEngine(modelDirectory: installedCache)
try await engine.ensureLoaded()
let result = try await engine.transcribe(samples: mono16kSamples)
print(result.text)

downloadedBundle is the extracted profile directory; installedCache is a fresh app-owned directory. Audio must be mono 16 kHz PCM. Keep the engine loaded between recordings and keep the Orukeet cache separate from Parakeet's. Use the explicit local loader rather than FluidAudio's NVIDIA model downloader.

Full conversion and integration guide · Source PR #6 · TapTalk

Measurements and scope

On one M5 Max, greedy decoding reduced paired batch latency by 9.5% relative to the Core ML baseline. Baseline and greedy returned identical text on 128 FLEURS recordings across eight languages. Core ML WER was 7.64%, compared with 8.55% for Parakeet Core ML and 7.29% for the uncompressed Orukeet source. Raw evidence and methodology.

The baseline uses 6-bit LUT/FP16 weights; the greedy profile removes unused top-64 joint calculations. It is unsuitable for decoding paths that require those outputs. The separate iOS experimental precision profile is described above.

These historical macOS measurements do not qualify other devices, older OS versions or power use. The current full-corpus INT8/LUT6 evaluation is documented separately above. FluidAudio's default sliding-window TDT path buffers roughly 13 seconds before first text. This is not the separate Parakeet EOU 120M live-typing model.

License and attribution

Orukeet modified weights: CC BY-SA 4.0, retaining NVIDIA Parakeet attribution. Core ML graphs derive from Fluid Inference's pinned Parakeet conversion. FluidAudio runtime: Apache-2.0. Conversion and integration code: MIT. The bundles retain LICENSE-WEIGHTS, NOTICE.md, and COREML-NOTICE.txt.

All weights derive from the r3 checkpoint with SHA-256 031c8ddab4845aeced904a7cde8e8aa57993b2e344716cf83a545b079c473b56. Baseline graph metadata retains the upstream names for graph identity; use bundle.json to identify the weights as Orukeet.