An audiobook you can talk to.
Your own textbooks and papers, read aloud, on your own computer. Tap your earbuds, ask out loud, and it answers from what you have heard so far and won't spoil what's ahead, then keeps reading.
No account, no subscription; by default nothing leaves your computer.
What it does
- Your own books and papers. PDFs, EPUBs, Word documents, slides, notes, or an arXiv id. Back matter nobody wants narrated, like references and the index, is skipped.
- Tap to ask. Tap your earbuds and it pauses and listens. The answer is grounded at the spot you were listening to, so "wait, what did it just say?" works, and it won't spoil what you haven't heard yet.
- Math spoken like a tutor. Formulas are read the way a tutor would say them ("the expectation over x of..."), not as symbol soup, and figures are described aloud. Built for dense STEM books.
- Questions that come back. After each chapter, an optional spoken quiz for active recall. The questions you asked come back on a schedule: yapbook re-asks them by voice, grades your spoken answer against the chapter text, and tracks what you have mastered.
- Whole-book memory. Ask "how does this relate to chapter 2?" or "what did it say earlier about X?". The answer draws on everything you have already heard, not on what you haven't yet.
- Your own voice as the narrator. Recorded once, on your computer, never uploaded: read four sentences, hear a paragraph rendered in that voice, then decide. Answers and every spoken system line stay in the default voice.
- A local web page. The library, player and settings live on a page served from your own computer: drag a book in, press play, and a Now Reading pane shows the sentence, equation or figure the earbuds are on.
See it
Measured
Measured during 0.2.0 development with the default fully local setup, on one machine: a Windows 11 laptop with an Intel Core Ultra 9 285HX, 64 GB RAM and an NVIDIA GeForce RTX 5090 Laptop GPU (24 GB), with the language, vision, embedding and speech-recognition models on the GPU and the narration voice and document parser on the CPU. Everything also runs on a CPU alone; a GPU just makes it snappier, so expect longer times without one.
| What | Result | How it was counted |
|---|---|---|
| Spoilers | 3% of verdicts, vs 85% for the same local model answering from the whole book with no position gate | Verdicts flagged as leaking not-yet-heard material, on questions whose answer lies past the playhead (36 questions, 72 verdicts). |
| Answers on heard material | Two thirds fully correct, a quarter partial, none wrong; about 1 in 12 refused when it could have been answered | Judge labels on questions whose answer was already heard (24 questions, 48 verdicts); wrong refusals counted over heard and mixed questions (72 verdicts). |
| "Not covered yet" honesty | 88% vs 0% | Verdicts where the answer says the book has not reached that yet, on questions whose answer is still ahead (24 questions, 48 verdicts). |
| Tap to first answer audio | Median 1.9 s, p90 2.8 s | From the end of your spoken question (voice-activity endpoint) to the first answer sentence synthesized as audio; 10 questions across 6 books, 3 timed runs each, the prompt prefix prewarmed at the previous segment boundary as in a live session, the question audio synthesized and fed through the real endpointer, nothing played. |
The first three rows are run 3 of a benchmark with 60 questions at known listening positions in 6 books; yapbook and the ungated comparison were answered by the same local 4B model and graded blind by two LLM judges, all three runs published. Caveats: the judges are language models, n is small, the latency run shared the GPU with another application at full utilization throughout, and the one spoiler yapbook still produces is the model answering a BatchNorm question from its own memory of that famous paper, documented item by item in the benchmark. Since run 3 the answer prompt names the book's author when it is known; a same-day check against a no-author control found the spoiler count unchanged (2 of 72 verdicts), 4 more "not covered yet" verdicts on unread material and 1 more wrong refusal (9 of 72 against 8), one run per arm, so the rows above stay run 3's.
The grounding benchmark is published with the code, all three runs and their verdicts.
Install
One line. It installs uv if you do not have it, installs yapbook in its own environment on Python 3.12, runs setup (asking before every install, pull or download, with the sizes printed first), then doctor. No admin rights.
Fully local: narration, transcription, the question-answering brain, figure descriptions and whole-book memory all run on your computer; about 10 to 12 GB of models (Ollama, Kokoro, Whisper and the document parser).
Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://raw.githubusercontent.com/Abinesh-L/yapbook/main/install.ps1 | iex"
macOS / Linux
curl -fsSL https://raw.githubusercontent.com/Abinesh-L/yapbook/main/install.sh | sh
Prefer an API key and no local brain? Set YAPBOOK_INSTALL_PROFILE to api before the one-liner. That profile downloads about 2.5 GB (the speech and PDF models plus a small embedder); narration, transcription and whole-book memory still run on your computer, and your spoken questions and the heard-so-far text go to the provider you choose. Figures keep caption placeholders until you add the local vision model later.
No book at hand? yapbook demo installs a bundled public-domain sample (four talks by William James, 1899) and starts reading.
Windows first, macOS next. Bluetooth earbuds are optional; a laptop mic is enough.
Until 0.2.0 is tagged (planned for Tuesday, September 29) the installer picks up 0.1.1, which has the read-aloud, tap-to-ask loop but not most of what this page describes: the whole-book memory, the "not covered yet" answers, the question book, the Now Reading pane, voice bookmarks and your own voice arrive with 0.2.0.
About
yapbook is a hobby project by Abinesh Lingeswaran, built for listening to dense textbooks while his hands are busy. Questions and feedback are welcome by email.