Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

AI Transcription and Retrieval System

A University of South Alabama capstone project developed for a publicly traded healthcare technology organization to address the challenge of making large volumes of information across multiple content formats easier to process, search, and retrieve.

The project focused on transforming audio, video, images, PDFs, and web-based content into searchable, timestamped outputs that could support contextual information retrieval and question answering. The system combined transcription, speaker diarization, visual-content interpretation, retrieval-augmented generation, and structured report generation.

Confidentiality Notice: Certain implementation details, source code, project materials, screenshots, and deliverables are intentionally omitted due to confidentiality obligations.

Business Problem

The organization needed a more efficient way to work with large amounts of information distributed across different file and media formats. Manually reviewing lengthy recordings, documents, images, and other content can make it difficult to quickly locate relevant information or use that information as context for downstream AI-assisted workflows.

The project explored how AI-based processing and retrieval could transform multimodal content into structured, searchable information for faster querying and review.

Project Solution

The team developed an AI-powered multimodal processing application that could ingest supported content, generate searchable and timestamped outputs, and make processed information available for retrieval-augmented question answering.

The application incorporated speech recognition, speaker diarization, visual analysis, document processing, retrieval-augmented generation, and report generation to support a unified information-processing workflow.

My Contributions

  • Collaborated on the design and development of the application
  • Contributed to audio and video transcription workflows
  • Worked with speaker diarization and timestamped outputs
  • Contributed to image and visual-content interpretation
  • Supported retrieval-augmented question answering
  • Worked with multimodal content including audio, video, images, PDFs, and URLs
  • Contributed to HTML, PDF, and DOCX report exports
  • Participated in testing, troubleshooting, and project documentation

Technologies

  • Python
  • Whisper ASR
  • pyannote.audio
  • LLaVA
  • FFmpeg
  • Retrieval-Augmented Generation (RAG)
  • LangChain
  • ChromaDB
  • Hugging Face
  • PyTorch
  • Ollama
  • Jinja2
  • ReportLab
  • python-docx

Core Capabilities

Transcription

Audio and video content could be processed into timestamped text using automatic speech recognition.

Speaker Diarization

Speaker diarization was used to help distinguish between different speakers within supported audio and video content.

Visual Content Analysis

Image-processing capabilities were incorporated to generate useful descriptions and contextual information from visual content.

Retrieval-Augmented Question Answering

Processed content could be indexed and retrieved to support grounded question answering based on uploaded or extracted information.

Report Generation

Processed results could be exported into multiple document formats, including:

  • HTML
  • PDF
  • DOCX

Skills Demonstrated

  • Python application development
  • Artificial intelligence integration
  • Natural language processing
  • Speech recognition
  • Speaker diarization
  • Multimodal AI workflows
  • Retrieval-augmented generation
  • Vector-based information retrieval
  • Document processing
  • Report generation
  • Debugging and troubleshooting
  • Collaborative software development
  • Technical documentation

Project Background

This project was completed as a university capstone involving the development of an AI-powered multimodal processing application.

Because portions of the project are subject to confidentiality obligations, this public overview intentionally excludes source code, internal materials, client information, proprietary workflows, and other restricted project artifacts.

Security and Confidentiality

This repository contains only a high-level description of the project.

It does not contain:

  • Proprietary source code
  • Client or company data
  • Internal documentation
  • Confidential screenshots
  • Credentials or API keys
  • Proprietary system architecture
  • Restricted project deliverables

Author

Todd Stringfellow

B.S. Information Technology
Digital Forensics Concentration
Minor in Computer Information Systems
University of South Alabama