An automated tool for the full cycle of Binotel call processing: downloading, transcription (local or cloud), multi-level AI analysis, and professional report generation.
- Asynchronous Pipeline (AsyncIO): All logic is built on non-blocking operations. The script can simultaneously download audio, transcribe via cloud services, and analyze data through AI, accelerating the processing of large archives by dozens of times.
- SQLite Caching & State Management: Replaces the legacy
state.jsonwith a robust SQLite database. It tracks every stage of the pipeline:- AI Analysis Caching: If power fails or a session is interrupted, the script won't waste money/tokens on AI for already processed calls.
- Atomicity: Data is written instantly and reliably.
- Token Bucket Rate Limiting: Intelligent control over API request frequency (RPM/TPM). This ensures maximum efficiency of Gemini, OpenAI, and Binotel limits, avoiding "429 Too Many Requests" errors.
- Multi-Period Analysis: Support for processing multiple date ranges simultaneously (e.g.,
2025-11-01:2025-12-31, 2026-03-01:2026-04-30) to compare seasonal performance. - Call Retrieval: Automatic history requests via Binotel API with adaptive chunking (if >2000 calls are found per period) and automatic deduplication.
- Smart Filtering: Automatically ignores SMS, technical records, and calls with zero duration.
- Multi-Provider Transcription:
- Local (WhisperX): Uses Faster-Whisper for high speed. Features precise diarization ('Speaker 1/2') and is optimized for 6GB VRAM GPUs (RTX 4050/3050).
- OpenAI Whisper API: Stable cloud transcription with automatic file splitting if they exceed 25MB.
- Deepgram Nova-2: The fastest cloud model with excellent support for multiple languages.
- Professional AI Analysis:
- Support for cutting-edge models: Gemini 1.5 Pro/Flash, Gemma 2, GPT-4o, Claude 3.5 Sonnet.
- Local AI: Free analysis capability via Ollama (Gemma 2, Llama 3).
- Security: Automated phone number and PII masking before sending data to cloud providers.
- Sales Script Generation: Based on the analysis of all calls, AI creates an ideal conversation script, considering typical manager mistakes and successful practices.
- Flexible Reporting:
- HTML: Interactive report with a dark theme, convenient call cards, and analysis costs.
- PDF: Professional format for printing or sending to management (full Cyrillic support).
- Cost Calculation: Detailed expense tracking based on tokens (
utils/models.json).
git clone https://github.com/ksalab/binotel-calls.git
cd binotel-calls- Python: Strictly 3.13. Using newer versions may lead to compatibility issues with WhisperX libraries.
- FFmpeg: Required for audio conversion and processing.
sudo add-apt-repository ppa:deadsnakes/ppa
sudo apt update
sudo apt install python3.13 python3.13-venv python3.13-dev ffmpeg cmake libcairo2-dev libxt-devpython3.13 -m venv .venv
source .venv/bin/activate # for Linux/macOS
# .\.venv\Scripts\activate # for Windowspip install -r requirements.txtCreate a .env file based on .env.example. Key parameters:
ANALYSIS_PERIODS: Dates separated by commas, e.g.,2025-01-01:2025-01-31.AI_RPM_LIMIT: AI requests per minute (default is15).TRANSCRIPTION_PROVIDER:local(GPU required),openai, ordeepgram.WHISPER_BATCH_SIZE:4for 6GB VRAM GPUs,8-16for more powerful ones.AI_PROVIDER:gemini,openai,ollama(for local AI).
# Full run (download + transcribe + analyze)
python get_calls.py --mode fetch
# Analysis of existing transcripts only (no Binotel API calls)
python get_calls.py --mode analysis
# System Health Check (keys, network, GPU, disk space)
python get_calls.py --check# General statistics and costs
python inspect_db.py --stats
# List last 20 calls (status of download, text, and analysis)
python inspect_db.py --list 20
# Detailed JSON analysis of a specific call
python inspect_db.py --details [CALL_ID]The project uses a sophisticated logging system with Trace IDs to make debugging easier in concurrent environments:
- Trace IDs: Every log entry is tagged with its context.
[TG-12345]: Logs related to a specific Telegram user.[CALL-f9a2b3]: Logs related to a specific call record processing.[SYSTEM]: General system events.
- Trace ID Middleware: In the Telegram bot, a middleware automatically extracts the user ID from each update and injects it into the logging context, ensuring all actions triggered by that user are traceable.
- Structured Logging: Set
LOG_FORMAT=jsonin your.envto enable machine-readable logs inlog/app.json.log, suitable for ELK/Grafana stacks. - Sticky Progress Bars: Console logs use
tqdm.writeto prevent log messages from breaking the progress bars.
The script automatically cleans the system based on .env parameters:
KEEP_RECORDS_DAYS=14: Deletes audio and transcripts older than 2 weeks.KEEP_LOGS_DAYS=30: Deletes old reports and log files.- Automated model cache cleanup (
.cache/torch,.cache/huggingface) to save space.
You can edit templates in the prompts/{lang}/ folder. Available variables:
{{DATA}}: Call metadata (duration, type, number).{{TRANSCRIPT}}: Full text of the conversation.{{INDIVIDUAL_ANALYSES}}: Aggregated data of all calls for the final report.{{RECOMMENDATIONS}}: Strategic tips for the sales script.
This happens due to a torch library update. Run:
python -c "import torch; import lightning; p = '.venv/lib/python3.13/site-packages/whisperx/assets/pytorch_model.bin'; c = torch.load(p, map_location='cpu', weights_only=False); c['pytorch-lightning_version'] = lightning.__version__; torch.save(c, p); print('Success')"- Verify that
HF_TOKENis correct. - Ensure you have accepted the terms for the
pyannote/segmentationandpyannote/speaker-diarizationmodels on Hugging Face.
Install DejaVu fonts: sudo apt install fonts-dejavu.
This project is licensed under the MIT License.