Skip to content

Repository files navigation

Binotel Calls Downloader & Transcriber

Читати українською

An automated tool for the full cycle of Binotel call processing: downloading, transcription (local or cloud), multi-level AI analysis, and professional report generation.

Key Features

  • Asynchronous Pipeline (AsyncIO): All logic is built on non-blocking operations. The script can simultaneously download audio, transcribe via cloud services, and analyze data through AI, accelerating the processing of large archives by dozens of times.
  • SQLite Caching & State Management: Replaces the legacy state.json with a robust SQLite database. It tracks every stage of the pipeline:
    • AI Analysis Caching: If power fails or a session is interrupted, the script won't waste money/tokens on AI for already processed calls.
    • Atomicity: Data is written instantly and reliably.
  • Token Bucket Rate Limiting: Intelligent control over API request frequency (RPM/TPM). This ensures maximum efficiency of Gemini, OpenAI, and Binotel limits, avoiding "429 Too Many Requests" errors.
  • Multi-Period Analysis: Support for processing multiple date ranges simultaneously (e.g., 2025-11-01:2025-12-31, 2026-03-01:2026-04-30) to compare seasonal performance.
  • Call Retrieval: Automatic history requests via Binotel API with adaptive chunking (if >2000 calls are found per period) and automatic deduplication.
  • Smart Filtering: Automatically ignores SMS, technical records, and calls with zero duration.
  • Multi-Provider Transcription:
    • Local (WhisperX): Uses Faster-Whisper for high speed. Features precise diarization ('Speaker 1/2') and is optimized for 6GB VRAM GPUs (RTX 4050/3050).
    • OpenAI Whisper API: Stable cloud transcription with automatic file splitting if they exceed 25MB.
    • Deepgram Nova-2: The fastest cloud model with excellent support for multiple languages.
  • Professional AI Analysis:
    • Support for cutting-edge models: Gemini 1.5 Pro/Flash, Gemma 2, GPT-4o, Claude 3.5 Sonnet.
    • Local AI: Free analysis capability via Ollama (Gemma 2, Llama 3).
    • Security: Automated phone number and PII masking before sending data to cloud providers.
  • Sales Script Generation: Based on the analysis of all calls, AI creates an ideal conversation script, considering typical manager mistakes and successful practices.
  • Flexible Reporting:
    • HTML: Interactive report with a dark theme, convenient call cards, and analysis costs.
    • PDF: Professional format for printing or sending to management (full Cyrillic support).
    • Cost Calculation: Detailed expense tracking based on tokens (utils/models.json).

Installation

1. Clone the repository

git clone https://github.com/ksalab/binotel-calls.git
cd binotel-calls

2. System Requirements & Dependencies

  • Python: Strictly 3.13. Using newer versions may lead to compatibility issues with WhisperX libraries.
  • FFmpeg: Required for audio conversion and processing.

Installation (Ubuntu/Debian)

sudo add-apt-repository ppa:deadsnakes/ppa
sudo apt update
sudo apt install python3.13 python3.13-venv python3.13-dev ffmpeg cmake libcairo2-dev libxt-dev

3. Create Virtual Environment

python3.13 -m venv .venv
source .venv/bin/activate  # for Linux/macOS
# .\.venv\Scripts\activate  # for Windows

4. Install Packages

pip install -r requirements.txt

Configuration (.env)

Create a .env file based on .env.example. Key parameters:

  • ANALYSIS_PERIODS: Dates separated by commas, e.g., 2025-01-01:2025-01-31.
  • AI_RPM_LIMIT: AI requests per minute (default is 15).
  • TRANSCRIPTION_PROVIDER: local (GPU required), openai, or deepgram.
  • WHISPER_BATCH_SIZE: 4 for 6GB VRAM GPUs, 8-16 for more powerful ones.
  • AI_PROVIDER: gemini, openai, ollama (for local AI).

Usage

1. Main Script

# Full run (download + transcribe + analyze)
python get_calls.py --mode fetch

# Analysis of existing transcripts only (no Binotel API calls)
python get_calls.py --mode analysis

# System Health Check (keys, network, GPU, disk space)
python get_calls.py --check

2. Database Utility

# General statistics and costs
python inspect_db.py --stats

# List last 20 calls (status of download, text, and analysis)
python inspect_db.py --list 20

# Detailed JSON analysis of a specific call
python inspect_db.py --details [CALL_ID]

📊 Observability & Debugging

The project uses a sophisticated logging system with Trace IDs to make debugging easier in concurrent environments:

  • Trace IDs: Every log entry is tagged with its context.
    • [TG-12345]: Logs related to a specific Telegram user.
    • [CALL-f9a2b3]: Logs related to a specific call record processing.
    • [SYSTEM]: General system events.
  • Trace ID Middleware: In the Telegram bot, a middleware automatically extracts the user ID from each update and injects it into the logging context, ensuring all actions triggered by that user are traceable.
  • Structured Logging: Set LOG_FORMAT=json in your .env to enable machine-readable logs in log/app.json.log, suitable for ELK/Grafana stacks.
  • Sticky Progress Bars: Console logs use tqdm.write to prevent log messages from breaking the progress bars.

Maintenance

The script automatically cleans the system based on .env parameters:

  • KEEP_RECORDS_DAYS=14: Deletes audio and transcripts older than 2 weeks.
  • KEEP_LOGS_DAYS=30: Deletes old reports and log files.
  • Automated model cache cleanup (.cache/torch, .cache/huggingface) to save space.

Prompts & Variables

You can edit templates in the prompts/{lang}/ folder. Available variables:

  • {{DATA}}: Call metadata (duration, type, number).
  • {{TRANSCRIPT}}: Full text of the conversation.
  • {{INDIVIDUAL_ANALYSES}}: Aggregated data of all calls for the final report.
  • {{RECOMMENDATIONS}}: Strategic tips for the sales script.

Troubleshooting

1. "Lightning automatically upgraded..." error

This happens due to a torch library update. Run:

python -c "import torch; import lightning; p = '.venv/lib/python3.13/site-packages/whisperx/assets/pytorch_model.bin'; c = torch.load(p, map_location='cpu', weights_only=False); c['pytorch-lightning_version'] = lightning.__version__; torch.save(c, p); print('Success')"

2. "Diarization finished: 0 unique speakers found"

  • Verify that HF_TOKEN is correct.
  • Ensure you have accepted the terms for the pyannote/segmentation and pyannote/speaker-diarization models on Hugging Face.

3. Cyrillic squares in PDF

Install DejaVu fonts: sudo apt install fonts-dejavu.

License

This project is licensed under the MIT License.

About

Automated Binotel call downloader, Ukrainian-language transcriber (WhisperX), and AI-powered business analytics engine with cost-tracking.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages