Cisco iconCiscoSep 24, 2026 ~7 min source read

Cisco Expands LLM Security Leaderboard to Rank Text, Image, and Audio Risks for Agentic Models

The Cisco LLM Security Leaderboard now evaluates multimodal attack surfaces — text, images, and audio — across 136 models so organizations can match model weaknesses to the specific modalities their agents will encounter.

LLM Security Leaderboard Now Ranks Every Modality Where Agent Risk Lives: Text, Image, and Audio

Share this story

Send the public story page.

Useful takeaways from this story.

Leaderboard expanded: 102 new evaluations since June 2026 bring the total to 136 models, tested without added guardrails to provide a consistent baseline.

Attack types tested: Evaluations include prompt injection, jailbreaks, single-turn and multi-turn text attacks, and comparable single-attempt image and audio attacks.

The useful part

When autocomplete results are available use up and down arrows to review and enter to select Downloads Certifications Design Guides Training Community Careers Cisco Blogs Cisco Blogs Artificial Intelligence. When we launched the Cisco LLM Security Leaderboard earlier this year, the goal was simple: give organizations clear, tested data on how models hold up against attacks, so they know the risks before they deploy one. That matters because AI models are increasingly built into products such as agents that read email, browse the web, and take actions on a person's behalf.

How it works

  • As always, we test models in their base configuration without additional guardrails, so scores reflect a consistent baseline for layering on additional security protections.
  • If a model only gets evaluated on text, that risk may not show up until it becomes a real incident.
  • Those differences show up directly in how a model resists attack on one modality versus another.
  • Image and audio attacks are tested the same way as single-turn text attacks (one attempt, one message), using the same attack and harm categories as its text score, so their resistance is comparable across...
  • Each model's overall Combined Score is now an average across every format it was evaluated on, and a new modality switch lets you isolate scores for text, image, or audio on their own.

What to take from it

Integration with AI Supply Chain Provenance Explorer Provenance matters because a model's weaknesses often aren't unique to that model. If two models share lineage, a vulnerability discovered in one can be present in the other, and stopping an investigation at the model currently deployed can miss where a problem actually originated or where else it might surface. That risk also varies by deployment: a model wired into a browsing agent is exposed on different inputs (or modalities such as text, images, and audio) than one only answering questions in a chat window, so where a specific model is weak matters as much as where it's strong.

Example or evidence

  • Multimodal results are now live Until now, the leaderboard measured text-based attacks two ways: single-turn, where one harmful message is sent straight to the model, and multi-turn, a longer back-and-forth...
  • Each of those labs takes a different approach to building and training multimodal capability, whether that's how image data flows into the LLM backbone, how much safety alignment goes into a vision or audio...
  • A model that can be manipulated could be turned against the person using it.
  • 102 new evaluations across modalities since June 2026 The LLM Security Leaderboard is one of the most comprehensive model security leaderboards.

Details worth keeping

--> LLM Security Leaderboard Now Ranks Every Modality. AI LLM Security Leaderboard Now Ranks Every Modality. Text, Image, and Audio Amy Chang Nicholas Conley Artificial Intelligence.

Related coverage

  • Knowbe4: Agent Risk Manager, KnowBe4's AI agent security product in the KnowBe4 Platform, integrates with the Claude Compliance API to help organizations monitor Claude activity and detect risks in real time.
  • Marktechpost: RSA launched Agent ID at The AI Conference in San Francisco.
  • Thehackernews: Autonomous security agents are getting good at finding bugs.
  • Cisco: Discover how Cisco security solutions and the four pillars of agentic security protect organizations against autonomous AI risks and evolving threats.

More context around this story.

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise
Marktechpost iconMarktechpostSep 30, 2026

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

RSA launched Agent ID at The AI Conference in San Francisco. It's an agentic identity security platform for finance, government, healthcare, and critical infrastructure. It has 3 modules. Discover finds sanctioned and shadow agents and MCP servers. Secure is an inline gateway that checks every tool call against policy.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app