πŸŽ™

Golos Bot: Telegram voice messages transcribed to text

in-house Telegram bot: voice notes, video notes, audio and video to text via Gemini

A Telegram bot that turns voice messages into text: voice notes, video notes, audio and video are transcribed verbatim, cleaned up or as a summary. Google Gemini, a pool of API keys, FFmpeg, a transcript cache; audio is never stored.

In production Β· in-house product
3 modes
verbatim, cleaned up, summary; switched with a button after transcription
180 minutes
default maximum recording length, adjustable in settings
30 days
transcript cache lifetime; the audio file itself is deleted right after processing

The task

Voice messages are convenient for the sender and inconvenient for the recipient. You cannot skim them, search them, quote them, or read them where sound is not an option: on the subway, in a meeting, at night next to a sleeping child. A five-minute voice message costs exactly five minutes of attention even when only two lines of it matter.

What was needed was a voice-to-text bot that works by simply forwarding a message. No sign-up on third-party sites, no uploading files in a browser, and honest behaviour: if there is no speech in the recording, the bot must say so rather than make something up.

The solution

Golos Bot is an in-house Telegram bot for transcribing voice messages. Implemented in version 2.2.0:

  • Any audio source. Voice messages (ogg), video notes (the audio track is used), audio files in mp3, m4a, ogg and wav, video (a clip from a call, a lecture), and documents when audio was sent as a file.
  • Three modes. Verbatim: every word as spoken, for agreements. Cleaned up: no fillers or repetition, split into paragraphs, for everyday use. Summary: key points and structure.
  • Redo differently. The mode is switched with a button under the finished transcript; there is no need to resend the recording.
  • Reply language. Russian, English or the original language.
  • As a file. A long transcript is delivered as .txt instead of being cut into pieces.
  • Cache. The same recording in the same mode comes back instantly, with no API cost.
  • Silence. If there is no speech, the bot reports that instead of inventing text.
  • Gemini key pool. Automatic switching when a quota is exhausted and automatic fallback to a backup model.
  • Limits. Transcriptions per hour and maximum recording length; the hourly limit can be changed from the admin panel on the fly. A premium list without the hourly limit; admins are always unlimited.
  • Admin panel. Statistics, analytics, keys, users, bans, broadcasts, settings, logs, maintenance. When a new user appears, admins get a card with the ID and username, once.

How it works

Gemini audio transcription in a single request

Google Gemini is a multimodal model: an AI model that understands sound directly, with no separate "recognise speech, then process the text" stage. The bot sends the audio together with a prompt (an instruction) that states the mode and language, and receives finished text. Fewer links in the chain means less loss of meaning. The model's reasoning depth and response timeout are configurable.

ffmpeg: checking and compressing to mono MP3

Before anything goes to the model, ffprobe (a utility from the FFmpeg toolkit) reads the recording's duration and checks that it contains sound. If the file is over 20 MB, FFmpeg recompresses it to mono MP3 at 32 kbps: that is enough for speech recognition, and the size drops several times over. The pydub library is deliberately not used: it has not been updated since 2021 and imports the audioop module, which was removed from Python 3.13.

Gemini key pool and quotas

The free Gemini quota is limited, so the bot keeps several keys. When a key hits its quota, it goes on cooldown (60 minutes by default) and the request goes through the next key. The user notices nothing; the admin gets a notification. The admin panel shows the state of each key and has a button to lift a cooldown manually. If the main model returns a 404 (renamed or withdrawn), the bot quietly switches to the backup model.

Transcript cache

The result is stored in SQLite under a "recording + mode" key. Sending the same recording in the same mode again returns the cached text instantly. The cache lives for a limited time (30 days by default); the audio file itself is deleted right after processing. The database also holds users and their settings, a transcription log, key usage and the settings that can be changed from the admin panel.

Administration and accounting

The admin panel opens with /admin and works on buttons under the message. Its header shows the bot version, so you can see what is actually deployed on the server. When a new user first appears, every admin receives a card: name, the ID as a separate block (copied with a tap), a link to the username and the total number of users in the database. The bot's logic is covered by unit checks and routing checks in which Telegram and Gemini are replaced with mocks, so no network is needed to run them.

Results

  • A voice message, video note, audio file or video becomes text: searchable, quotable and readable without sound.
  • The mode can be changed after transcription without resending the recording.
  • Repeats are served from the cache without calling the API.
  • One key running out of quota does not stop the service: a key pool and a backup model.
  • Audio is not stored; the transcript lives for a limited time.
  • The bot runs in Docker with a healthcheck and log rotation; the version is written to the log at startup and sent to admins in a startup message.

Technologies and why

  • Python 3.14 and aiogram 3: an asynchronous Telegram bot driven by inline buttons.
  • Google Gemini (google-genai): audio transcription in a single request; main and backup models are configurable.
  • FFmpeg and ffprobe: duration, sound detection, compression of long recordings.
  • SQLite (aiosqlite): users, log, transcript cache, key usage, settings.
  • Docker: python:3.14-slim image, non-root user, healthcheck, Moscow time zone.

Status

In-house product, in production. The current version is 2.2.0 from 20 September 2026: it added the "Made in DevUnit Lab" signature with a link to the site. Version 2.1.0 (15 August 2026) marked the move to a separate repository and the split of the single-file bot into a package of modules.

Limitations: a 20 MB download cap through the cloud Bot API, private chats only, no queue when every key is exhausted. Planned: a "Copy text" button, a summary as the first message, group support, quota monitoring, cost accounting, export to .docx and .srt.

Questions about this project

How can I read a voice message without listening to it?
Forward the voice message to the bot. It replies with text in the mode you chose: verbatim, cleaned up (no fillers or repetition) or a summary. A long transcript arrives as a .txt file rather than in chunks.
Which recordings does the bot understand?
Voice messages, video notes (the audio track is used), audio files in mp3, m4a, ogg and wav, video, and documents when the audio was sent as a file. Reply language: Russian, English or the original language.
How does audio transcription with Gemini work?
The audio and an instruction (mode and language) go to the multimodal Google Gemini model in a single request. There is no separate speech-to-text step followed by text processing. Before sending, ffprobe checks the duration and whether there is sound, and FFmpeg recompresses files over 20 MB to mono MP3.
What happens when a Gemini key runs out of quota?
Keys work as a pool. An exhausted key goes on cooldown (60 minutes by default), the request goes through the next key, and the admin gets a notification. If the main model returns a 404, the bot switches to the backup model on its own.
Are my recordings stored?
The audio file is deleted right after processing. The transcript stays in a cache for a limited time, 30 days by default: the same recording in the same mode comes back instantly with no API cost.
Are there any limits?
Telegram's cloud Bot API does not hand bots files over 20 MB. Gemini safety filters may reject a recording on a sensitive topic, and noise or several people talking at once lowers quality. If there is no speech in the recording, the bot says so instead of inventing text.

Need something similar?

Tell us about the task β€” we'll show how we solved it and estimate the scope.