1.2.0
Extract text from Word DOCX files.
License evidence: MIT — installed notice
ไม่มีเวลาอ่าน? แปลงเอกสารเป็นไฟล์เสียง แล้วเก็บไว้ฟังเมื่อสะดวก
รุ่นทดสอบแบบมีรหัสเชิญ · บริการเสียงทำงานเมื่อเครื่องของผู้ดูแลออนไลน์ รับงานครั้งละหนึ่งเอกสาร กรุณาทดสอบด้วยเอกสารที่ไม่เป็นความลับ และดาวน์โหลดเสียงเก็บไว้ก่อนครบ 24 ชั่วโมง
Use case · Document listening
Sometimes you want to take in a document but don’t have time to sit down and read it. Upload your Thai, English or Chinese document, or paste text, and turn it into an MP3 to listen to when convenient. Listening instead of reading is the main purpose; reviewing notes and rehearsing are optional uses.
Invited beta · Access by numeric code
Validation snapshot: 11 September 2026. Speech quality still needs human listening and acceptance.
Current scope: up to 200,000 characters and a 4 MiB upload. Long documents become listening parts, with individual MP3s, one combined MP3 or a ZIP. Failed parts can be retried without repeating completed ones. PDF and complex Word layouts are outside this version.
12 September update: Mandarin and chapter downloads passed local tests. The backend suite now has 30 passing tests, including a failed middle part, service restart, retry without regenerating completed parts, and ZIP contents matching the MP3 parts. Earlier benchmark results below remain the 11 September snapshot.
The long-document run used 19,984 characters and produced 1,196 seconds of audio in 117.66 seconds on the test machine. These are observed local results, not a speed guarantee.
Sonnet assisted with implementation and a separate QA review. Codex inspected and corrected the code, integrated it and ran the tests. Opus reviewed the candidate and returned a local pass. These were explicit model calls, not an autonomous Big Crew run.
The prototype generates speech locally without a paid TTS API. Hosting costs and pronunciation accuracy are not guaranteed.
แชร์เครื่องมือที่ใช้จริงเพื่อให้คนอื่นศึกษาต่อ: หน้าที่ของแต่ละตัว เวอร์ชันที่ทดสอบ และลิงก์ต้นทาง ขอบคุณผู้พัฒนาและชุมชนที่ทำให้โปรเจกต์นี้เป็นไปได้
Selected libraries, models and tools, not a complete dependency audit. The Thai runtime’s license remains unresolved; its model’s MIT declaration does not cover every package automatically.
1.2.0
Extract text from Word DOCX files.
License evidence: MIT — installed notice
5.3.7
Find Thai word boundaries before splitting long passages.
License evidence: Apache-2.0 — project declaration
2026.9.10
Handle Unicode graphemes and decorative emoji without splitting Thai tone marks.
License evidence: Apache-2.0 AND CNRI-Python — installed metadata
revision 6fa5f03a26f8
Generate the Thai voice in the local prototype.
License evidence: MIT declared for the model; separate runtime unresolved
0.0.7
Load the Thai voice and turn Thai text into speech samples.
License evidence: UNRESOLVED — installed wheel has no license evidence; do not assume model license applies
code 0.9.4; model revision f3ff3571791e
Generate English and Mandarin speech locally; Mandarin voice zf_xiaobei.
License evidence: Apache-2.0 — code metadata and model card
1.30.0
Execute the Thai ONNX model.
License evidence: MIT — installed notice
2.14.0
Execute Kokoro’s English neural speech model.
License evidence: See upstream LICENSE and bundled third-party notices
3.8.16 / 3.8.0
Provide English linguistic processing used by the speech pipeline.
License evidence: Model MIT and LICENSES_SOURCES observed; review each package’s notices
0.14.0
Write generated speech samples to intermediate audio files.
License evidence: See upstream license and libsndfile dependency notices
8.1.1 local binary
Normalize and join audio sections, then encode the final MP3.
License evidence: GPL-3.0-or-later for the tested enabled-GPL build; build-dependent
0.141.1 / 0.52.4
Serve the private upload, audio-job and download API.
License evidence: FastAPI MIT observed; review Uvicorn notices separately
existing website stack
Build the upload editor, audio player, use-case card and report.
License evidence: See each project’s LICENSE and the website lockfile
test tooling
Test backend behavior and real browser upload, playback and downloads.
License evidence: See each project’s LICENSE; these are development tools
0.9.4
Convert Chinese text to the phonemes used by Kokoro.
License evidence: Apache-2.0 — installed metadata
0.42.1 / 0.55.0
Segment Chinese words and derive Mandarin pronunciations.
License evidence: MIT — installed metadata for both packages
0.5.24
Convert Arabic numerals to Chinese spoken forms.
License evidence: MIT — installed metadata
0.1.7
Convert traditional characters to simplified for speech while preserving the original preview.
License evidence: Apache License — installed metadata