Tilawi FastConformer (Quran recitation, on-device)
The speech model behind the Tilawi Quran app: it runs on the phone, so recitations never leave the device. Given a recitation, it produces a CTC transcript that @tilawi/quran-asr matches against the fixed Quran text to find the recited verse, or to give word-by-word memorization feedback.
- Fine-tuned from NVIDIA's
stt_ar_fastconformer_hybrid_large_pcd_v1.0(CTC branch) on the author's own recitation recordings. - Exported to a single ONNX graph with the preprocessor inside, so the input is raw audio. Mixed quantization (int4 MatMul, int8 Conv/LayerNorm): 88 MB.
- Runs with
onnxruntime-react-native(iOS/Android) oronnxruntime-node.
Links: tilawi.ai (the app) · github.com/Tilawi (all open source) · npm @tilawi/quran-asr, @tilawi/react-native-quran-asr, @tilawi/expo-pcm-recorder
The model never outputs or alters Quran text. It transcribes the user's recitation; the matcher compares that transcript to the canonical text.
Files
| File | What it is |
|---|---|
fastconformer_full_mixed.onnx |
The model. Input audio_signal float32 [1, N] (16 kHz mono PCM) and length int64 [1]; output log-probabilities [1, T, 1025]. |
vocab.json |
CTC token id to SentencePiece piece (1,025 entries; blank id 1024). |
quran_ctc_tokens.json |
Each verse's text as model token ids, for CTC re-ranking. |
quran.json |
6,236 verse records (surah, ayah, text_uthmani, text_clean, surah names). Tanzil text; see TANZIL-NOTICE.txt. |
SHA256SUMS |
Checksums of the four files above. |
Accuracy
Verse identification through @tilawi/quran-asr. The numbers come from 600 verses (200 per reciter, sampled across the whole Quran, seed 1) from EveryAyah; real live tests, with people reciting into phones, were done as well:
| Reciter | Correct verse | No match | Wrong verse |
|---|---|---|---|
| Mohamed Siddiq al-Minshawi (murattal) | 200 / 200 | 0 | 0 |
| Mahmoud Khalil al-Husary | 200 / 200 | 0 | 0 |
| Mishary Rashid Alafasy | 199 / 200 | 1 | 0 |
"Correct" includes 12 verses whose text is word-for-word identical to another verse (e.g. 55:13), which audio cannot tell apart. The harness and per-clip results are in the quran-asr repo.
Limitations. The table uses studio recordings because they can be scored automatically; the model has also been tested with live recitation recorded on phones. It isn't perfect: noise, very short verses or unclear recitation can lead to a missed word or no match, and the matcher prefers no answer over a wrong one. It is trained for Quran recitation (Hafs), not general Arabic speech. It's free and open source in the hope that it helps.
Use
In your own project: npm install @tilawi/quran-asr onnxruntime-node, then see the usage example. To try it from the repo:
git clone https://github.com/Tilawi/quran-asr && cd quran-asr
npm install && npm run build && npm run fetch-assets
node examples/transcribe.mjs recitation.mp3
Credits and license
- Model weights: CC-BY-4.0. Base model: NVIDIA,
stt_ar_fastconformer_hybrid_large_pcd_v1.0, CC-BY-4.0. Fine-tuning and export: Muhammed Durakovic. - Quran text in
quran.json: Tanzil Project, CC-BY-3.0, redistributed verbatim; the full notice is inTANZIL-NOTICE.txt. The token table is derived from that text.