Skip to content

Repository files navigation

CyberIntelBot

A threat intelligence assistant over MITRE ATT&CK, and an incident analyser that reads logs.

Retrieval-augmented generation: MITRE ATT&CK techniques are embedded into a Pinecone index, a question retrieves the relevant ones, and a Groq-hosted model answers from what was retrieved and cites the technique id so every claim can be checked against attack.mitre.org.

B.Tech Final Year Project, SASTRA Deemed University, May 2025.


Two modes, one chat box

Ask a question. "How does APT29 get initial access?" retrieves the matching techniques and answers from them. Technique ids are cited, and any id the model produces that was not in the retrieved context is flagged on screen as unverified.

Drop logs. Attach log files or an incident report to the composer and it works out what the attacker did, maps each step onto ATT&CK, and returns containment and remediation steps taken from MITRE's own mitigation text.

Accepted: .log, .txt, .csv, .json, .jsonl, .xml, .tsv, .md and .pdf. Format is detected from the content, not the extension, so a CSV named .log still parses with its column names intact. The attachment button is a folder picker, because log evidence usually arrives as a directory per host.


How the incident analysis works

Raw log lines are never embedded. A syslog line is mostly timestamp, host, PID and address, and embedding that buries the behaviour under the noise: two unrelated lines from the same host at the same minute come out more similar than the same attack seen on two different machines.

So a line is taken apart first, and then two independent layers look at it:

layer what it is what it is good at what it misses
rule pack ~55 patterns in incident.py, each mapped to a real ATT&CK id precision, and it can prove itself by quoting the line that matched anything nobody wrote a rule for
retrieval the RAG index, queried with cleaned behaviour phrases finding techniques nobody anticipated it is a similarity score, and it will be wrong

Every finding says which layer produced it. A rule hit is confirmed and comes with the literal log line; a retrieval hit is possible and comes with the phrase that made it look similar. Retrieval only runs on behaviour the rules did not already explain, so the two arms cannot corroborate each other in a circle.

A retrieval candidate also needs agreement from two separate behaviours before it is reported at all. Without that threshold this arm returns the same handful of broad techniques for almost any input, which is the failure the keyword evaluation below already documented at the question level.

The dual use problem, and why the tool reports a verdict

Run the rule pack against a log of one ordinary working day, backups and a software deployment and some helpdesk troubleshooting, and it reports fifteen techniques, four of them high severity. Every one is genuinely present: mapping a C$ share is T1021.002, registering a scheduled task is T1053.005, running whoami is T1033. The rules are right and the report is useless.

ATT&CK techniques are overwhelmingly dual use, because attackers work by doing ordinary things for hostile reasons. Treating "this technique occurred" as "this is an attack" is the mistake, and it is the one that gets a detection tool switched off.

So techniques that administrators perform routinely are marked, kept out of the verdict, and presented as context rather than as findings. The verdict counts only high severity detections that are not dual use:

                            decisive   dual use   verdict
ransomware_endpoint.log           13         13   incident indicated
linux_ssh_intrusion.log            7          9   incident indicated
web_app_compromise.jsonl           8          6   incident indicated
benign_admin_day.log               0         15   no incident indicated

This will be wrong in both directions on real traffic. An attacker who uses nothing but dual use techniques, which is the entire idea behind living off the land, scores zero on it. It is a triage aid, not a verdict to act on.


Measurements

Everything below was produced by running the code, not estimated.

All of the incident measurements are fully offline: no index, no API key, no network, under a second.

HOSTILE LOGS, mean over 3 incidents: recall 1.000, precision 1.000
  60 of 60 planted techniques found
BENIGN LOG, 39 lines of ordinary administration:
  15 techniques reported, 4 of them high severity
  0 decisive after the dual use filter

Real third-party logs, python evaluate_incident.py --real

18,000 lines of Loghub production logs from nine systems: Linux, Apache, OpenSSH, Hadoop, Zookeeper, Windows, Spark, Mac and a proxy. Real data, collected years before this tool existed.

REAL LOGS, 9 sources, 17,999 lines nobody wrote for this tool
  techniques raised beyond known attack traffic: 3
  of those, decisive (would trigger a verdict):  2
  rate: 0.17 per 1,000 lines

It caught T1110 on the OpenSSH set, which is 28 days from a real internet-facing server and genuinely full of brute force. That is a true positive on traffic nobody staged, and it is the most meaningful single result in this project.

It also found a real bug in its own rule pack on first contact. T1033 had a pattern \bid\s*$ for the Unix id command. On real logs it fired on "mapred.tip.id is deprecated ... use mapreduce.task.id" and on a macOS line ending "effectiveBundleID": any line ending in the word id matched. Deleted, for the same reason as T1112. That is the entire argument for third-party data, and it paid out within a minute of first use.

Of the 3 remaining, none is clearly wrong: a burst of PAM authentication failures on a public Linux host, TeamViewer on a proxy log, and an accepted SSH login. The last two are dual use and correctly do not carry a verdict.

Synthetic incidents, python evaluate_incident.py --rules

Read the second half of that, not the first. The logs are synthetic, the ground truth was written by the same person as the rules, and on logs where every line is hostile recall cannot fall. The 1.000 is a regression test, not a performance claim. The benign count is the only number there that can punish a loosely written rule, and it is the one that caught the dual use problem.

The single biggest gap: no false positive rate against real traffic. Getting one needs real logs.

Retrieval, python evaluate.py

Measured against 10 hand written questions on a live index of 607 techniques.

SUMMARY over 10 questions at k=10
                     recall@k   precision@k       MRR
  naive                 0.933         0.340     0.715
  keyword               0.850         0.276     0.650
  improvement           -8.9%        -18.8%     -9.1%

LLM keyword extraction makes retrieval worse on all three metrics. The cause was diagnosed rather than guessed: N keywords each contribute their own top 10 into a pool truncated back to 10, and generic terms poison it. Keyword selectivity, the fraction of a keyword's own nearest neighbours that literally contain it:

adversaries 100%    credentials 84%    persistence 58%    DNS 54%
LSASS 6%            APT29 2%           spear phishing attachments 0%

"adversaries" literally matches everything it retrieves, so it carries zero discriminating power while still consuming slots. Its remaining justification is multi-turn follow up handling, which the metrics do not measure.

Indexing, python benchmark_index.py

Ground truth for free: every technique carries description, which is indexed, and x_mitre_detection, which never is. 579 techniques have both, so the detection text becomes a query whose answer is known and whose words the index has never seen. No hand labelling, and the labels cannot be wrong.

INDEXING STRATEGY, 579 held out queries, k=10
strategy                          recall@10       MRR
baseline: whole, truncated            0.693     0.474      579 vectors
chunked: name + chunks                0.699     0.494    1,508 vectors
improvement                           +1.0%     +4.3%

Splitting by whether that technique's description was actually truncated shows the interesting part: chunking did not help the group it exists to help. The recall gain came entirely from the untruncated group, so the measured win is the name prefix, not the chunking. MITRE descriptions are front loaded, and 44% of the discarded tail's vocabulary already appears in the kept portion.

The sister project, a resume search tool, got +125% from the identical fix, because resumes are back loaded: skills and projects sit at the end. Same bug, same fix, 125x difference in payoff, entirely explained by document structure.

Latency

keyword extraction   mean 0.25s      retrieval   mean 1.75s
generation           mean 0.94s      TOTAL       mean 2.94s   median 2.86s

A first attempt measured 10.34s and the per question times climbed steadily. That was Groq's burst rate limiter, not the pipeline: spacing the questions 20s apart dropped generation to 0.94s. Benchmarking back to back against a rate limited API measures the rate limiter.


Hybrid retrieval

Dense retrieval fused with BM25 by reciprocal rank fusion, in hybrid.py, on by default.

The reasoning comes straight out of the measurement above. LITERAL_MATCH_BONUS adds a flat 0.15 when a keyword appears anywhere in a document: a term weighting scheme with one term and one weight, unable to tell "APT29", which appears in 2% of what it retrieves, from "adversaries", which appears in all of it. BM25 is the same idea with inverse document frequency doing that separation properly, and exact identifiers are precisely the case dense embeddings handle worst.

Fusion is by rank, not score: BM25 scores and cosine similarities have no common scale, and any weighted sum needs a constant tuned on one corpus that will not transfer. RRF needs none.

It is not assumed to help. On the sister project four of five standard retrieval upgrades made the main task worse, including the one predicted to be the biggest single win. Set USE_HYBRID = False for the old behaviour, and quote no number from this that did not come out of the evaluation harness.


Setup

pip install -r requirements.txt

Credentials come from the environment, never from source. Copy .env.example to .env next to the script:

PINECONE_API_KEY=your_key_here
GROQ_API_KEY=your_key_here

.env is gitignored. Shell variables also work but only last for that window.

streamlit run cyberintelbot.py

First run populates the index: 607 live techniques, 2,185 chunks, a few minutes. Later runs skip it.

Running the checks

python test_retrieval.py                  83 checks, offline, mocked
python test_incident.py                  645 checks, offline
python evaluate_incident.py --rules      the incident numbers, offline
python evaluate.py                       keyword vs naive, needs a live index
python evaluate_incident.py --retrieval  rules vs RAG on the same logs, live index
python benchmark_index.py                indexing strategies, needs a live index

Try it without an index of your own: the four labelled logs in sample_incidents/ drive evaluate_incident.py --rules and test_incident.py with no credentials at all.


Notes on models

Groq retires hosted models without warning. This project pinned llama-3.3-70b-versatile and every Llama chat model has since been withdrawn from this account, returning 404 model_not_found. Retrieval was unaffected, so it failed as "the language model call failed" while the index worked perfectly. There is now a fallback list and the UI names whichever model answered, so a silent downgrade is impossible.

Two traps worth keeping:

  • gpt-oss models reason before replying, spending completion tokens on thinking that never reaches content. At the max_tokens=50 the keyword extractor used to run with, and at 128, they return an empty string with no error. Measured: 50 empty, 128 empty, 512 works. Swapping the model alone would have left the bot silently broken.
  • Cloudflare blocks plain urllib against the Groq API with error 1010, and it looks exactly like a 403 auth failure on every model. Use httpx.

List what a key can actually call:

python -c "import httpx,os; print(httpx.get('https://api.groq.com/openai/v1/models', headers={'Authorization':'Bearer '+os.environ['GROQ_API_KEY']}).json())"

Layout

cyberintelbot.py        the app: retrieval, generation, incident wiring, UI
incident.py             log parsing, IOC extraction, the rule pack, triage
hybrid.py               BM25, reciprocal rank fusion, the chunk corpus
evaluate.py             keyword retrieval vs naive, needs a live index
evaluate_incident.py    the incident pipeline against labelled logs
benchmark_index.py      indexing strategies, using MITRE detection text
test_retrieval.py       83 offline checks, Pinecone and Groq mocked
test_incident.py        645 offline checks
sample_incidents/       four labelled logs, three hostile and one benign

Known limitations, volunteer these

  • The incident ground truth is synthetic and was written alongside the rules. Recall of 1.000 measures that the rules match what they were written to match. The real-log run is the honest half of the evaluation.
  • 18,000 lines of real logs is a start, not a baseline. It is nine systems over a few days, mostly Linux and application logs, so it exercises perhaps a third of the rule pack and barely touches the Windows half.
  • The dual use list was read off that one benign log. The right idea fitted to far too little data. A real deployment would derive it per estate from a few weeks of its own logs.
  • The rule pack is English, Windows and Linux, and text. No EVTX parsing, no cloud audit formats, no network packet capture.
  • Keyword extraction lowers retrieval quality and is kept for multi-turn follow ups, which nothing here measures.
  • Line order is taken as chronological. Timestamps are not parsed out of every format, so a shuffled or merged multi-host log will have a wrong timeline while every individual finding stays correct.
  • Hybrid retrieval is unmeasured on this corpus. The reasoning is sound and the numbers are not in yet.

About

Maps raw incident logs onto MITRE ATT&CK and returns containment and remediation steps. Deterministic detection rules plus hybrid BM25 and vector retrieval, with every finding labelled by which layer found it.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages