Let’s Hear It
Local text-to-speech reader for macOS
Let’s Hear It is a personal macOS app for listening to English EPUBs and text-based PDFs. You import a book, pick a voice, and listen while the current sentence is highlighted. Book files, extracted text, speech generation, playback, and progress stay on the Mac.
Architecture
The app has two processes. A native SwiftUI app owns the interface, library, and audio playback. A Python worker runs the speech models through MLX Audio on Apple Silicon. The app starts the worker and talks to it over stdin and stdout; no HTTP server is involved.
Components
- Importer. Reads the EPUB container, package document, spine, and HTML directly, keeping spine order and table-of-contents labels. PDFKit extracts text-based PDFs with page provenance and basic cleanup, such as repeated headers. Scanned PDFs are rejected with an explanation.
- Library. Imported books are copied into Application Support under a content hash, so duplicates reuse one copy and books survive moving the original file. SQLite holds metadata, navigation, progress, and audio-cache records; books, covers, and audio are files.
- Speech worker. Runs Kokoro (the default) or the experimental Chatterbox Turbo through MLX Audio. Model revisions are pinned; after the voice download, the worker runs in offline mode.
- Player. AVFoundation plays one audio file per sentence with pitch-preserving speed from 0.75× to 2×. Highlighting follows audio that has actually played, not generation progress.
- Updater. Sparkle checks a signed appcast on the public releases repository daily and verifies each download before installing.
Listening loop
- The reader splits the current chapter into sentence units.
- The app requests the current sentence and a bounded look-ahead from the worker. Each request carries an ID.
- The worker writes the audio atomically and reports its path and measured duration.
- The player queues the file and advances the highlight when that sentence finishes playing.
- Seeking or changing voice cancels obsolete requests. Late results with stale IDs are discarded.
Generated audio is cached by text, model revision, voice, and text-processing version, so a different voice or model never reuses old audio. The cache has a 2 GB default limit and evicts least-recently-used audio without touching books or progress.
Releases
The application source is private. A release script builds an Apple Silicon DMG, signs the update feed, and publishes the DMG, checksums, and appcast.xml to the publicletshearit-updates repository. It stays public so the app can fetch updates without credentials.
Scope
Requires an Apple Silicon Mac on macOS 15 or later. Scanned PDFs, OCR, voice cloning, audio export, and phone sync are not included.