Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🔦 Project Lumos

AI-Powered Security Posture Management (AI-SPM) for Shadow AI Discovery

Project Lumos uses unsupervised machine learning (K-Means Clustering) to detect unauthorized AI tool usage ("Shadow AI") by analyzing network flow metadata. It identifies employees using unapproved LLM providers — such as personal ChatGPT accounts, Claude, or AI coding assistants — without deep packet inspection.


Problem Statement

Organizations face a Declaration-Observation Gap: employees declare they only use approved AI tools, but network observations reveal massive egress to unauthorized LLM providers. This creates:

  • Data Exfiltration Risk - Sensitive IP sent to external models for "summarization"
  • Compliance Drift - Usage of models that violate EU AI Act or TRAIGA requirements
  • Shadow IT Blindspot - Security teams can't manage what they can't see

Technical Approach

Component Implementation
Data Source NetFlow / VPC Flow Logs (metadata only, no DPI)
Model K-Means Clustering (Unsupervised Learning)
Features 22 engineered features: packet size distributions, inter-arrival times, destination entropy, burst scores, streaming ratios
Classification Clusters compared against AI Provider Fingerprint database
Dashboard Bootstrap 5 dark-mode admin panel (Adminator-inspired)

How It Works

  1. Data Ingestion - Synthetic NetFlow records simulate enterprise traffic (web, SaaS, video, file transfer, DNS) mixed with AI provider traffic
  2. Feature Engineering - Extracts 22 features that capture the AI traffic "signature": burst outbound (prompts) + streaming inbound (token responses)
  3. K-Means Clustering - Groups flows by behavior using Elbow Method + Silhouette Score for optimal k
  4. Fingerprint Matching - Maps clusters to known AI providers via IP prefix and hostname matching
  5. Risk Assessment - Compliance scoring against EU AI Act and TRAIGA frameworks

Quick Start

# Install dependencies
pip install -r requirements.txt

# Run the ML pipeline
python main.py

# Serve the dashboard
python -m http.server 8080

# Open in browser
# http://localhost:8080/dashboard/index.html

Project Structure

project-lumos/
|-- config/
|   |-- ai_provider_fingerprints.json   # Known AI provider signatures
|-- src/
|   |-- __init__.py
|   |-- data_generator.py              # Synthetic NetFlow generator
|   |-- feature_engineering.py         # 22-feature extraction pipeline
|   |-- clustering.py                  # K-Means + Elbow + Silhouette
|   |-- classifier.py                  # Shadow AI classifier
|-- dashboard/
|   |-- index.html                     # Adminator-style Bootstrap 5 UI
|   |-- style.css                      # Dark/light theme system
|   |-- app.js                         # Chart.js visualizations
|-- output/                            # Generated by pipeline
|   |-- lumos_results.json             # Dashboard data payload
|   |-- raw_flows.csv                  # Raw synthetic flows
|-- main.py                            # Pipeline orchestrator
|-- requirements.txt
|-- README.md

Dashboard Sections

Section Description
Overview Summary cards, 24h traffic timeline, cluster donut, provider bar chart, detection metrics
Cluster Analysis PCA scatter plot, cluster legend, elbow method chart, silhouette scores
Shadow AI Providers Per-provider cards with flow counts, data volume, risk level
Compliance EU AI Act and TRAIGA status panels, compliance detail matrix
ML Model Pipeline info, feature list, confusion matrix

Performance Results

Metric Value
Precision 69.3%
Recall 99.7%
F1 Score 81.7%
Optimal K 8
Features 22

High recall (99.7%) means almost no Shadow AI goes undetected. The precision/FP tradeoff reflects that some SaaS traffic shares similar patterns with AI traffic - a known challenge in unsupervised network analysis.

AI Provider Fingerprint Database

The system tracks these providers:

Provider Risk Level Category
OpenAI (ChatGPT) Critical General Purpose LLM
Anthropic (Claude) High General Purpose LLM
Google AI Studio (Gemini) High General Purpose LLM
Cursor AI Critical AI Coding Assistant
Hugging Face Medium Model Hosting
Replicate Medium Model Hosting
Microsoft Copilot Approved Enterprise AI

Tech Stack

  • Python 3.11+ with NumPy, pandas, scikit-learn
  • Bootstrap 5.3 (Adminator-inspired dark mode)
  • Chart.js 4 for interactive visualizations
  • K-Means unsupervised clustering algorithm

License

MIT

About

AI-Powered Security Posture Management (AI-SPM) for Shadow AI Discovery. Uses K-Means clustering on network flow metadata to detect unauthorized AI tool usage.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages