The TWIML AI Podcast
The TWIML AI Podcast

CTIBench: Evaluating LLMs in Cyber Threat Intelligence with Nidhi Rastogi - #729

Today, we're joined by Nidhi Rastogi, assistant professor at Rochester Institute of Technology to discuss Cyber Threat Intelligence (CTI), focusing on her recent project CTIBench—a benchmark for evaluating LLMs on real-world CTI tasks. Nidhi explains the evolution of AI in cybersecurity, from r

Featured Speakers

Nidhi Rastogi Guest

Topics Discussed

Episode Summary

Executive Summary: Nidhi Rastogi discusses how LLMs are reshaping cyber threat intelligence by adding context, speed, and retrieval over traditional rules and ML, while also introducing risks like hallucinations and stale knowledge. She explains CTI Bench, a practical benchmark built from real-world security sources to test models on knowledge, attribution, vulnerability mapping, and severity scoring, and outlines future work on mitigation, drift, and explainability.

Main Topics: LLMs change cybersecurity from pattern detection to contextual intelligence (Priority: 5/5): Rastogi contrasts rule-based and classical ML approaches with LLMs that can ingest threat reports, logs, and other modalities to provide faster, more context-aware security reasoning. Cyber threat intelligence (CTI) as a multimodal security discipline (Priority: 5/5): CTI is framed as the aggregation of reports, logs, indicators of compromise, and attack techniques to detect, attribute, and anticipate threats across an organization. RAG and freshness as critical for security use cases (Priority: 4/5): Because LLMs have knowledge cutoff dates, retrieval/fine-tuning approaches are important to keep models current with new malware, CVEs, and emerging attack patterns. CTI Bench as a practical benchmark for evaluating LLMs (Priority: 5/5): The lab created CTI Bench to measure whether general-purpose and security-specific LLMs can answer cybersecurity questions accurately using tasks modeled after a threat analyst's daily work. Model limitations: hallucinations, blind spots, and uneven performance (Priority: 5/5): Even strong models can misattribute threat actors, confuse security concepts, or fail on complex questions, showing the need for human-in-the-loop oversight. Future directions: mitigation, explainability, and concept drift (Priority: 4/5): The team is moving beyond detection and benchmarking toward evaluating mitigation guidance, model confidence/explainability, and methods to detect and correct drift over time.

Key Arguments: LLMs help cybersecurity by turning fragmented security data into contextual answers much faster than a human analyst could. RAG and similar retrieval approaches are necessary because static model training cannot keep pace with rapidly changing threats and vulnerabilities. Cyber threat intelligence requires reasoning over heterogeneous sources, not just logs; LLMs can help unify reports, telemetry, and structured standards. A useful benchmark in cybersecurity must reflect real analyst tasks, not generic academic probes. CTI Bench reveals that even top models perform unevenly: strong on straightforward knowledge tasks, weaker on complex or recent security questions. Hallucinations are especially dangerous in cybersecurity because a confident but wrong recommendation can mislead analysts and organizations. Open-source models such as Llama can be surprisingly capable, but performance still depends heavily on model size and training corpus. Specialized models like SecGemini suggest domain fine-tuning can materially improve performance on cybersecurity tasks. Future value lies in pairing LLMs with explainability and mitigation validation so analysts know when to trust, verify, or override the model.

Data Points: RIT faculty experience: 4 years - Rastogi has been at Rochester Institute of Technology working at the intersection of AI and cybersecurity. CTI Bench question set: 2,500 questions - Generated with ChatGPT-4 and then heavily reviewed and corrected by humans. Threat reports used: 50 reports - Standard reports from MITRE were used to create and validate attack-pattern tasks. Model sizes mentioned: 7B, 8B, 70B parameters - Llama models of varying sizes were tested to compare cybersecurity performance. Performance frequency for stronger models: 3 out of 4 times - Advanced models often performed well on tasks, while smaller models struggled more often. Benchmark task count: 4 to 5 task types - CTI Bench grouped questions into multiple cybersecurity reasoning and knowledge categories. Publication timing: mid-2024 - The benchmark evaluation described was submitted to NeurIPS around that period.

Pivotal Quotes: "We'd identified a space where there is no such existing benchmarking model which can tell whether this specific language model is capable of giving good responses or accurate responses on cybersecurity specific tasks." — Nidhi Rastogi: Explaining why CTI Bench was created "What an analyst would have taken like a couple of hours to analyze and to comprehend and then associate with a threat pattern they have identified in their network are able to do that in a couple of seconds." — Nidhi Rastogi: Describing the practical value of LLMs in CTI "It is very, very important. So it was good to know that even the best of the best models are able to fail sometimes." — Nidhi Rastogi: Discussing why benchmarks should expose blind spots

Implications: Security teams should treat LLMs as assistants, not authorities: use them for speed and synthesis, but verify outputs against current sources and human expertise. Benchmarks like CTI Bench will likely shape safer, more specialized cyber AI systems.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast