A Retrieval-Augmented Generation (RAG) based AI assistant that transforms lecture videos into a searchable knowledge base and provides accurate, context-aware answers to students' questions. The system leverages FFmpeg for audio extraction, OpenAI Whisper for transcription, BGE-M3 for semantic embeddings, and an Ollama/OpenAI GPT model for response generation.
- 🎥 Convert lecture videos into audio using FFmpeg
- 🎙️ Generate timestamped transcripts using OpenAI Whisper
- 🧩 Split transcripts into semantic chunks
- 🧠 Generate vector embeddings using BGE-M3
- 📦 Store embeddings in a serialized Joblib knowledge base
- 🔍 Retrieve relevant lecture content using Cosine Similarity
- 🤖 Generate accurate, grounded responses using an OpenAI GPT model
RAG_BASED_AI/
│── jsons/ # Transcript JSON files
│── unused/ # Experimental scripts
│── .env # Environment variables
│── .gitignore # Git ignore rules
│── embeddings.joblib # Serialized vector knowledge base
│── mp3_to_json.py # Audio → JSON transcripts
│── preprocess_json.py # Generate embeddings
│── process_incoming.py # RAG inference pipeline
│── prompt.txt # Prompt template
│── README.md
│── requirements.txt
│── response.txt # Generated response (optional)
│── video_to_mp3.py # Video → Audio conversion
git clone https://github.com/aryanraj7791/RAG_based_AI_teaching_assistant.git
cd RAG_based_AI_teaching_assistantpip install -r requirements.txtCreate a .env file in the project root.
OPENAI_API_KEY=your_openai_api_key
Place all lecture videos inside the videos/ directory.
Convert lecture videos into MP3 files.
python video_to_mp3.pyTranscribe each MP3 file into a timestamped JSON transcript.
python mp3_to_json.pyGenerate semantic embeddings for every transcript chunk.
python preprocess_json.pyThis step:
- Reads transcript JSON files
- Splits transcripts into semantic chunks
- Generates vector embeddings using BGE-M3
- Stores embeddings and metadata in embeddings.joblib
Run the inference pipeline.
python process_incoming.pyThe application:
- Loads the vector knowledge base.
- Converts the user query into an embedding.
- Retrieves the most relevant transcript chunks using cosine similarity.
- Constructs a Retrieval-Augmented prompt.
- Sends the prompt to an OpenAI GPT model.
- Returns a personalized, context-aware response grounded in the lecture content.
Lecture Videos
│
▼
FFmpeg (Video → Audio)
│
▼
OpenAI Whisper (Speech-to-Text)
│
▼
JSON Transcripts
│
▼
Text Chunking
│
▼
BGE-M3 Embedding Model
│
▼
embeddings.joblib
│
──────────────────────────────────
│
Student Question
│
▼
Query Embedding
│
▼
Cosine Similarity Search
│
▼
Relevant Context Retrieval
│
▼
Prompt Construction
│
▼
OpenAI GPT Model
│
▼
Personalized Response
| Category | Technologies |
|---|---|
| Language | Python |
| Video Processing | FFmpeg |
| Speech-to-Text | OpenAI Whisper |
| Embedding Model | BGE-M3 |
| Large Language Model | OpenAI GPT |
| Data Processing | Pandas, NumPy |
| Machine Learning | Scikit-learn |
| Knowledge Base | Joblib |
| Similarity Search | Cosine Similarity |
| Environment Management | python-dotenv |
- 🗄️ Vector database integration (FAISS/Qdrant)
- 🌐 Web-based interface using Flask or Streamlit
- 📚 Multi-course knowledge base
- 🔗 Source citations for retrieved responses
Aryan Raj
Data Scientist | AI Engineer | Full Stack Developer
- 🌐 GitHub: https://github.com/aryanraj7791
- 💼 LinkedIn: https://www.linkedin.com/in/aryan-raj-79246b280/
- 📧 Email: aryanraj5371@gmail.com
⭐ If this project helped you, please star the repository!