A threat intelligence assistant over MITRE ATT&CK, and an incident analyser that reads logs.
Retrieval-augmented generation: MITRE ATT&CK techniques are embedded into a Pinecone index, a question retrieves the relevant ones, and a Groq-hosted model answers from what was retrieved and cites the technique id so every claim can be checked against attack.mitre.org.
B.Tech Final Year Project, SASTRA Deemed University, May 2025.
Ask a question. "How does APT29 get initial access?" retrieves the matching techniques and answers from them. Technique ids are cited, and any id the model produces that was not in the retrieved context is flagged on screen as unverified.
Drop logs. Attach log files or an incident report to the composer and it works out what the attacker did, maps each step onto ATT&CK, and returns containment and remediation steps taken from MITRE's own mitigation text.
Accepted: .log, .txt, .csv, .json, .jsonl, .xml, .tsv, .md and
.pdf. Format is detected from the content, not the extension, so a CSV named
.log still parses with its column names intact. The attachment button is a
folder picker, because log evidence usually arrives as a directory per host.
Raw log lines are never embedded. A syslog line is mostly timestamp, host, PID and address, and embedding that buries the behaviour under the noise: two unrelated lines from the same host at the same minute come out more similar than the same attack seen on two different machines.
So a line is taken apart first, and then two independent layers look at it:
| layer | what it is | what it is good at | what it misses |
|---|---|---|---|
| rule pack | ~55 patterns in incident.py, each mapped to a real ATT&CK id |
precision, and it can prove itself by quoting the line that matched | anything nobody wrote a rule for |
| retrieval | the RAG index, queried with cleaned behaviour phrases | finding techniques nobody anticipated | it is a similarity score, and it will be wrong |
Every finding says which layer produced it. A rule hit is confirmed and comes
with the literal log line; a retrieval hit is possible and comes with the
phrase that made it look similar. Retrieval only runs on behaviour the rules did
not already explain, so the two arms cannot corroborate each other in a
circle.
A retrieval candidate also needs agreement from two separate behaviours before it is reported at all. Without that threshold this arm returns the same handful of broad techniques for almost any input, which is the failure the keyword evaluation below already documented at the question level.
Run the rule pack against a log of one ordinary working day, backups and a
software deployment and some helpdesk troubleshooting, and it reports fifteen
techniques, four of them high severity. Every one is genuinely present:
mapping a C$ share is T1021.002, registering a scheduled task is T1053.005,
running whoami is T1033. The rules are right and the report is useless.
ATT&CK techniques are overwhelmingly dual use, because attackers work by doing ordinary things for hostile reasons. Treating "this technique occurred" as "this is an attack" is the mistake, and it is the one that gets a detection tool switched off.
So techniques that administrators perform routinely are marked, kept out of the verdict, and presented as context rather than as findings. The verdict counts only high severity detections that are not dual use:
decisive dual use verdict
ransomware_endpoint.log 13 13 incident indicated
linux_ssh_intrusion.log 7 9 incident indicated
web_app_compromise.jsonl 8 6 incident indicated
benign_admin_day.log 0 15 no incident indicated
This will be wrong in both directions on real traffic. An attacker who uses nothing but dual use techniques, which is the entire idea behind living off the land, scores zero on it. It is a triage aid, not a verdict to act on.
Everything below was produced by running the code, not estimated.
All of the incident measurements are fully offline: no index, no API key, no network, under a second.
HOSTILE LOGS, mean over 3 incidents: recall 1.000, precision 1.000
60 of 60 planted techniques found
BENIGN LOG, 39 lines of ordinary administration:
15 techniques reported, 4 of them high severity
0 decisive after the dual use filter
18,000 lines of Loghub production logs from nine systems: Linux, Apache, OpenSSH, Hadoop, Zookeeper, Windows, Spark, Mac and a proxy. Real data, collected years before this tool existed.
REAL LOGS, 9 sources, 17,999 lines nobody wrote for this tool
techniques raised beyond known attack traffic: 3
of those, decisive (would trigger a verdict): 2
rate: 0.17 per 1,000 lines
It caught T1110 on the OpenSSH set, which is 28 days from a real internet-facing server and genuinely full of brute force. That is a true positive on traffic nobody staged, and it is the most meaningful single result in this project.
It also found a real bug in its own rule pack on first contact. T1033 had a
pattern \bid\s*$ for the Unix id command. On real logs it fired on
"mapred.tip.id is deprecated ... use mapreduce.task.id" and on a macOS line
ending "effectiveBundleID": any line ending in the word id matched. Deleted, for
the same reason as T1112. That is the entire argument for third-party data, and
it paid out within a minute of first use.
Of the 3 remaining, none is clearly wrong: a burst of PAM authentication failures on a public Linux host, TeamViewer on a proxy log, and an accepted SSH login. The last two are dual use and correctly do not carry a verdict.
Read the second half of that, not the first. The logs are synthetic, the ground truth was written by the same person as the rules, and on logs where every line is hostile recall cannot fall. The 1.000 is a regression test, not a performance claim. The benign count is the only number there that can punish a loosely written rule, and it is the one that caught the dual use problem.
The single biggest gap: no false positive rate against real traffic. Getting one needs real logs.
Measured against 10 hand written questions on a live index of 607 techniques.
SUMMARY over 10 questions at k=10
recall@k precision@k MRR
naive 0.933 0.340 0.715
keyword 0.850 0.276 0.650
improvement -8.9% -18.8% -9.1%
LLM keyword extraction makes retrieval worse on all three metrics. The cause was diagnosed rather than guessed: N keywords each contribute their own top 10 into a pool truncated back to 10, and generic terms poison it. Keyword selectivity, the fraction of a keyword's own nearest neighbours that literally contain it:
adversaries 100% credentials 84% persistence 58% DNS 54%
LSASS 6% APT29 2% spear phishing attachments 0%
"adversaries" literally matches everything it retrieves, so it carries zero discriminating power while still consuming slots. Its remaining justification is multi-turn follow up handling, which the metrics do not measure.
Ground truth for free: every technique carries description, which is indexed,
and x_mitre_detection, which never is. 579 techniques have both, so the
detection text becomes a query whose answer is known and whose words the index
has never seen. No hand labelling, and the labels cannot be wrong.
INDEXING STRATEGY, 579 held out queries, k=10
strategy recall@10 MRR
baseline: whole, truncated 0.693 0.474 579 vectors
chunked: name + chunks 0.699 0.494 1,508 vectors
improvement +1.0% +4.3%
Splitting by whether that technique's description was actually truncated shows the interesting part: chunking did not help the group it exists to help. The recall gain came entirely from the untruncated group, so the measured win is the name prefix, not the chunking. MITRE descriptions are front loaded, and 44% of the discarded tail's vocabulary already appears in the kept portion.
The sister project, a resume search tool, got +125% from the identical fix, because resumes are back loaded: skills and projects sit at the end. Same bug, same fix, 125x difference in payoff, entirely explained by document structure.
keyword extraction mean 0.25s retrieval mean 1.75s
generation mean 0.94s TOTAL mean 2.94s median 2.86s
A first attempt measured 10.34s and the per question times climbed steadily. That was Groq's burst rate limiter, not the pipeline: spacing the questions 20s apart dropped generation to 0.94s. Benchmarking back to back against a rate limited API measures the rate limiter.
Dense retrieval fused with BM25 by reciprocal rank fusion, in hybrid.py, on by
default.
The reasoning comes straight out of the measurement above. LITERAL_MATCH_BONUS
adds a flat 0.15 when a keyword appears anywhere in a document: a term weighting
scheme with one term and one weight, unable to tell "APT29", which appears in 2%
of what it retrieves, from "adversaries", which appears in all of it. BM25 is
the same idea with inverse document frequency doing that separation properly,
and exact identifiers are precisely the case dense embeddings handle worst.
Fusion is by rank, not score: BM25 scores and cosine similarities have no common scale, and any weighted sum needs a constant tuned on one corpus that will not transfer. RRF needs none.
It is not assumed to help. On the sister project four of five standard
retrieval upgrades made the main task worse, including the one predicted to be
the biggest single win. Set USE_HYBRID = False for the old behaviour, and
quote no number from this that did not come out of the evaluation harness.
pip install -r requirements.txt
Credentials come from the environment, never from source. Copy .env.example to
.env next to the script:
PINECONE_API_KEY=your_key_here
GROQ_API_KEY=your_key_here
.env is gitignored. Shell variables also work but only last for that window.
streamlit run cyberintelbot.py
First run populates the index: 607 live techniques, 2,185 chunks, a few minutes. Later runs skip it.
python test_retrieval.py 83 checks, offline, mocked
python test_incident.py 645 checks, offline
python evaluate_incident.py --rules the incident numbers, offline
python evaluate.py keyword vs naive, needs a live index
python evaluate_incident.py --retrieval rules vs RAG on the same logs, live index
python benchmark_index.py indexing strategies, needs a live index
Try it without an index of your own: the four labelled logs in
sample_incidents/ drive evaluate_incident.py --rules and test_incident.py
with no credentials at all.
Groq retires hosted models without warning. This project pinned
llama-3.3-70b-versatile and every Llama chat model has since been withdrawn
from this account, returning 404 model_not_found. Retrieval was unaffected,
so it failed as "the language model call failed" while the index worked
perfectly. There is now a fallback list and the UI names whichever model
answered, so a silent downgrade is impossible.
Two traps worth keeping:
- gpt-oss models reason before replying, spending completion tokens on
thinking that never reaches
content. At themax_tokens=50the keyword extractor used to run with, and at 128, they return an empty string with no error. Measured: 50 empty, 128 empty, 512 works. Swapping the model alone would have left the bot silently broken. - Cloudflare blocks plain
urllibagainst the Groq API with error 1010, and it looks exactly like a 403 auth failure on every model. Usehttpx.
List what a key can actually call:
python -c "import httpx,os; print(httpx.get('https://api.groq.com/openai/v1/models', headers={'Authorization':'Bearer '+os.environ['GROQ_API_KEY']}).json())"
cyberintelbot.py the app: retrieval, generation, incident wiring, UI
incident.py log parsing, IOC extraction, the rule pack, triage
hybrid.py BM25, reciprocal rank fusion, the chunk corpus
evaluate.py keyword retrieval vs naive, needs a live index
evaluate_incident.py the incident pipeline against labelled logs
benchmark_index.py indexing strategies, using MITRE detection text
test_retrieval.py 83 offline checks, Pinecone and Groq mocked
test_incident.py 645 offline checks
sample_incidents/ four labelled logs, three hostile and one benign
- The incident ground truth is synthetic and was written alongside the rules. Recall of 1.000 measures that the rules match what they were written to match. The real-log run is the honest half of the evaluation.
- 18,000 lines of real logs is a start, not a baseline. It is nine systems over a few days, mostly Linux and application logs, so it exercises perhaps a third of the rule pack and barely touches the Windows half.
- The dual use list was read off that one benign log. The right idea fitted to far too little data. A real deployment would derive it per estate from a few weeks of its own logs.
- The rule pack is English, Windows and Linux, and text. No EVTX parsing, no cloud audit formats, no network packet capture.
- Keyword extraction lowers retrieval quality and is kept for multi-turn follow ups, which nothing here measures.
- Line order is taken as chronological. Timestamps are not parsed out of every format, so a shuffled or merged multi-host log will have a wrong timeline while every individual finding stays correct.
- Hybrid retrieval is unmeasured on this corpus. The reasoning is sound and the numbers are not in yet.