Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🔍 Copied Post Detector

A lightweight, database-free Streamlit app that flags whether a "Comparison Post" is Original, Paraphrased/Partial Match, or Plagiarised/Copied relative to a "Reference Post" — using classic NLP techniques (TF-IDF + Cosine Similarity, and Jaccard Similarity). No vector databases, no deep learning models, no external APIs.

✨ Features

  • Text or image input — type posts directly, or upload a screenshot and extract text via OCR (pytesseract).
  • Two similarity algorithms — TF-IDF Cosine Similarity and word-level Jaccard Similarity, averaged into a combined score.
  • Adjustable threshold — sidebar slider controls the cutoff for each verdict.
  • Color-coded verdict — ✅ Original (green), ⚠️ Paraphrased/Partial Match (yellow), 🚨 Plagiarised/Copied (red).
  • Visual dashboard — Plotly gauge chart, bar chart of scores vs. threshold, highlighted keyword overlap, and a demo "System Performance" panel.
  • Graceful edge-case handling — empty inputs, unreadable images, and OCR failures all show friendly messages instead of crashing.

🗂 Project Structure

copied-post-detector/
├── app.py             # Main Streamlit application
├── requirements.txt   # Python dependencies
├── packages.txt        # System package (Tesseract OCR) for Streamlit Cloud
├── .gitignore
└── README.md

🚀 Run Locally

  1. Clone/download the project, then create a virtual environment:

    python -m venv venv
    source venv/bin/activate      # Windows: venv\Scripts\activate
  2. Install Python dependencies:

    pip install -r requirements.txt
  3. Install the Tesseract OCR binary (required for the image-upload feature):

    • macOS: brew install tesseract
    • Ubuntu/Debian: sudo apt install tesseract-ocr
    • Windows: install from the UB-Mannheim Tesseract build and ensure it's on your PATH.

    The text-comparison features work without Tesseract installed — only the "Upload image" OCR option needs it.

  4. Run the app:

    streamlit run app.py

    It will open at http://localhost:8501.

☁️ Deploy for Free on Streamlit Community Cloud

  1. Push the project to GitHub:

    git init
    git add .
    git commit -m "Initial commit: Copied Post Detector"
    git branch -M main
    git remote add origin https://github.com/<your-username>/copied-post-detector.git
    git push -u origin main
  2. Go to share.streamlit.io and sign in with your GitHub account.

  3. Click "New app", then select:

    • Repository: <your-username>/copied-post-detector
    • Branch: main
    • Main file path: app.py
  4. Click "Deploy".

    • Streamlit Cloud automatically installs everything in requirements.txt.
    • It also reads packages.txt and installs tesseract-ocr at the system level, so the OCR upload feature works out of the box — no extra configuration needed.
  5. Your app will be live at a URL like: https://<your-username>-copied-post-detector-app-xxxxxx.streamlit.app

  6. Redeploying: any future git push to the connected branch automatically redeploys the app.

🧠 How the Scoring Works

  • TF-IDF Cosine Similarity — represents each post as a weighted vector of word importance, then measures the cosine of the angle between the two vectors (1 = identical direction, 0 = unrelated).
  • Jaccard Similarity — a simpler word-overlap measure: |shared words| ÷ |total unique words| (with common English stopwords filtered out).
  • Combined Score — the average of the two, compared against your chosen threshold:
    • Score ≥ threshold → 🚨 Plagiarised/Copied
    • Score ≥ 60% of threshold → ⚠️ Paraphrased/Partial Match
    • Otherwise → ✅ Original

⚠️ Notes & Limitations

  • This is a lightweight, educational similarity tool — not a legal or forensic plagiarism-detection system.
  • The "System Performance" tab shows illustrative/demo precision, recall, and F1 figures for dashboard presentation purposes; only the processing-time metric is measured live.
  • OCR accuracy depends on image quality/resolution — low-quality screenshots may extract partial or garbled text.

About

A lightweight, database-free Streamlit web application for detecting copied or plagiarised content from text and images using NLP techniques like TF-IDF and Jaccard Similarity.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages