Categories

AI Agents

How to use Shieldstral as a local guardrail for your agent

Shieldstral is the 3B safety classifier Mistral released in August 2026: it answers yes or no to a policy written in plain language. Converted to GGUF Q8_0 with llama.cpp and served on a CPU, it blocked no legitimate one among the 100 Spanish and English prompts I prepared, but let 8 of 17 injections through.

Artificial Intelligence

LLM red teaming: a practical playbook

LLM red teaming has gone from an esoteric activity to a mandatory practice. With the OWASP Agentic Top 10 and the CSA Agentic AI Red Teaming Guide converging on shared vocabulary, this is the operational playbook any team deploying agents needs to have.

Artificial Intelligence

LLM guardrails: frameworks and their real cost

Guardrails frameworks promise to filter language-model inputs and outputs to block data leaks, harmful content, or hallucinations. After evaluating four of the most popular ones in production, I cover what they actually do, what latency and billing cost they add, and when they pay off over simpler controls.

Artificial Intelligence

LLM agent security: the new class of threats

LLM agent security became a real incident category the moment assistants gained tool access. Agents that call tools expose a far wider attack surface than chatbots, and it now carries assigned CVEs, published audit reports, and its own OWASP Top 10. The spread of the Model Context Protocol and agent-driven corporate workflows built that surface in under 18 months.