An AI-powered story generation microservice that fetches real-time global trends and generates narratives using Microsoft's Phi-2 language model — served over both REST and gRPC protocols.
Features • Architecture • Quick Start • API Reference • Screenshots • Tech Stack
TrendStory is a production-style NLP microservice built as part of an advanced NLP course project. It bridges real-world trend data and generative AI — pulling live topics from Google News and YouTube, then generating contextual stories in three emotional tones using Microsoft Phi-2, a state-of-the-art 2.7B parameter language model.
The system is designed with dual API protocols (REST + gRPC), containerized with Docker, and includes a Gradio UI for interactive exploration.
Key Highlight for Recruiters: This project demonstrates end-to-end MLOps thinking — from real-time data ingestion and LLM inference to API design, gRPC communication, Docker deployment, and automated API testing.
| Feature | Description |
|---|---|
| Real-Time Trend Fetching | Pulls live trending topics from GNews API and YouTube Data API |
| AI Story Generation | Uses Microsoft Phi-2 (2.7B params) via HuggingFace Transformers |
| Multi-Tone Narratives | Generates stories in tragic, hopeful, or poetic tones |
| Dual API Protocols | REST (FastAPI) + gRPC — production microservice pattern |
| Text-to-Speech | Optional TTS playback of generated stories |
| Gradio UI | Interactive web interface for demos |
| Docker Support | Fully containerized, single-command deployment |
| API Testing Suite | Postman collection + Python benchmark script |
| Fallback Model | TinyLlama fallback for low-memory environments |
┌────────────────────────────────────────────────────────────┐
│ CLIENT LAYER │
│ Gradio UI │ Postman │ gRPC Client │
└──────────────┬──────────────────────┬──────────────────────┘
│ REST (HTTP/JSON) │ gRPC (Protobuf)
▼ ▼
┌────────────────────────────────────────────────────────────┐
│ API LAYER │
│ FastAPI REST Server │ gRPC Server (server.py) │
│ POST /generate_story │ StoryService.Generate() │
│ GET /trends │ │
└──────────────┬─────────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────┐
│ TREND INGESTION │
│ GNews API ──────────────────┐ │
│ YouTube Data API ───────────┼──► Topic Aggregator │
└──────────────────────────────┬────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────┐
│ LLM INFERENCE ENGINE │
│ Microsoft Phi-2 (HuggingFace Transformers) │
│ Fallback: TinyLlama (low-memory environments) │
│ Prompt: topic + tone → structured story narrative │
└────────────────────────────────────────────────────────────┘
| Gradio UI | Postman API Testing |
![]() |
![]() |
- Python 3.10+
- 8GB+ RAM (for Phi-2) or 4GB+ (TinyLlama fallback)
- GNews API Key & YouTube Data API Key
- Docker (optional but recommended)
git clone https://github.com/alihashim786/NLP-Agentic-Microservice-Platform.git
cd NLP-Agentic-Microservice-Platform
pip install -r requirements.txtcp .env.example .env
# Add your GNews API key and YouTube Data API keyuvicorn main:app --reload
# Server starts at http://127.0.0.1:8000docker build -t trendstory .
docker run -p 8000:8000 trendstorypython gradio_ui.py
# Interactive UI at http://127.0.0.1:7860Returns a merged list of trending topics from GNews and YouTube.
curl http://127.0.0.1:8000/trendsResponse:
{
"trends": [
"Floods in Lahore",
"Pakistan Elections 2025",
"Tech Layoffs Google",
"..."
]
}Generates an AI story for a given topic and tone.
curl -X POST http://127.0.0.1:8000/generate_story \
-H "Content-Type: application/json" \
-d '{"topic": "Floods in Lahore", "tone": "tragic"}'Payload:
{
"topic": "Floods in Lahore",
"tone": "tragic" // Options: "tragic" | "hopeful" | "poetic"
}Response:
{
"story": "Once upon a time, in the flood-stricken streets of Lahore, the river rose without warning..."
}Tone Examples:
| Tone | Description |
|---|---|
tragic |
Dark, sorrow-ful narrative highlighting human cost |
hopeful |
Optimistic story of resilience and recovery |
poetic |
Lyrical, metaphor-rich prose style |
The gRPC interface exposes the same story generation capability over Protocol Buffers for high-performance, low-latency clients.
# 1. Compile the .proto file
python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. trendstory.proto
# 2. Start gRPC server
python server.py
# 3. Call from gRPC client
python client.pyImport trendstory_postman_collection.json into Postman:
- Set
base_url = http://127.0.0.1:8000 - Includes tests for valid requests, missing fields, and invalid tones
python test_api.pyTests all API cases and measures response latency for performance benchmarking.
| Layer | Technology | Purpose |
|---|---|---|
| Language Model | Microsoft Phi-2 (2.7B) | Story generation |
| ML Framework | HuggingFace Transformers | LLM inference |
| REST API | FastAPI + Uvicorn | HTTP endpoints |
| RPC Protocol | gRPC + Protobuf | High-perf RPC interface |
| Data Sources | GNews API + YouTube Data API | Real-time trends |
| UI | Gradio | Interactive web demo |
| TTS | gTTS / pyttsx3 | Text-to-speech playback |
| Containerization | Docker | Deployment |
| API Testing | Postman | Endpoint validation |
| Fallback Model | TinyLlama | Low-memory environments |
nlp-agentic-microservice-platform/
├── main.py # FastAPI server (REST endpoints)
├── gradio_ui.py # Interactive Gradio interface
├── server.py # gRPC server
├── client.py # gRPC client
├── trendstory.proto # Protobuf service definition
├── test_api.py # Automated API tests + benchmarks
├── Dockerfile # Docker containerization
├── requirements.txt
├── trendstory_postman_collection.json # Postman test collection
├── i220583,i220554.ipynb # Development notebook
└── README.md
Why Phi-2? Microsoft's Phi-2 is a 2.7B parameter model that punches above its weight in reasoning and text generation tasks, making it a practical choice for story generation without requiring enterprise-grade GPU infrastructure.
Why REST + gRPC? The dual-protocol design demonstrates real-world microservice patterns. REST is ideal for web clients and external integrations; gRPC is preferred for internal service-to-service communication due to lower latency and binary protocol efficiency.
Why Docker? Containerization ensures reproducible environments across development, testing, and production — a core DevOps principle.
- Phi-2 model is ~7GB; a CUDA-capable GPU or 16GB+ RAM is recommended for fast inference
- gRPC client and server must be started separately
- API keys are required for live trend fetching (GNews + YouTube)
- No fine-tuning yet for regional language prompts (e.g., Roman Urdu)
- Fine-tune Phi-2 on regional news datasets
- Add streaming story generation (token-by-token)
- Implement rate limiting and API key auth
- Add story rating and feedback loop
- Kubernetes deployment manifests
- CI/CD pipeline with GitHub Actions
Built as an NLP course project at FAST-NUCES (2025).
- Hamza Jaffer — i220583
- [Co-author] — i220554
This project is licensed under the MIT License — see LICENSE for details.

