CYBERSECURITYTRACKER
TRACKING6,902 stories in this site build1,444 vulnerability news stories in this site build
Permanent story citation

Benchmarking the Agentic SOC: How we evaluate LLMs for security workflows

This page keeps the story as Cybersecurity Tracker first published it. If the tracker later corrects it, the correction appears below the original and never replaces it.

Back to newsStory 3746

As cited

Copy frozen at (site build).

ai security

Benchmarking the Agentic SOC: How we evaluate LLMs for security workflows

Elastic Security has built an evaluation framework to benchmark large language models (LLMs) used in agentic security operations centers (SOCs), moving beyond generic LLM leaderboards to measure real-world security task performance. The framework seeds realistic intrusions into a live deployment, runs multiple models through the same security tasks, and evaluates not just outputs but every tool call and parameter passed, grading the work rather than the writing. The evaluation focuses on seven capability categories including alert analysis, entity analytics, threat hunting, detection rules, workflow authoring, triggering workflows, and multi-step response chaining.

Why it matters: SOC teams evaluating LLM-driven automation need to validate that models reliably choose the correct tools, execute them in the right order, and ground their answers in actual tool outputs rather than fabricating plausible results, since a confident wrong answer in security triage is a missed intrusion.

Source published
First seen by Cybersecurity Tracker

Source attribution

Glossary