AR-03 / JUL 2026 / FULL STACK, SOLO
SUBTITLE LIVE — JA → ID
ffmpeg taps the PipeWire loopback and cuts roughly five-second chunks; a local speech model transcribes them; a local LLM translates, summarises on an interval, and explains cultural terms. The result streams to the browser over WebSocket as live subtitles.
The interesting work is around the pipeline: silence-based cutting so sentences are not sliced in half, a queue so chunks cannot collide, searchable history with transcript export, furigana and romaji, click-a-word lookup, Anki export, a transparent theatre overlay, and speaker filtering so only one voice gets translated. Eleven phases, about 2,200 lines, three models sharing 8 GB of VRAM.