NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
-
Updated
Aug 7, 2026 - Python
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
DeepTeam is a framework to red team LLMs and AI agents.
[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
A resource repository for machine unlearning in large language models
Agent trace and tool-use safety evaluation lab.
Decrypted Generative Model safety files for Apple Intelligence containing filters
Static security scanner for LLM agents — prompt injection, MCP config auditing, taint analysis. 51 rules mapped to OWASP Agentic Top 10 (2026). Works with LangChain, CrewAI, AutoGen.
[NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.
Papers about red teaming LLMs and Multimodal models.
Attack to induce LLMs within hallucinations
Open Source Reliability Harness: Make your agents follow rules. One line of code to enforce, trace, and improve.
Reading list for adversarial perspective and robustness in deep reinforcement learning.
Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses | 500+ Papers | Perception, Cognition, Planning, Interaction, Agentic System
Offline prompt-contract auditor for Hermes, Claude Code, Codex, OpenCode & OpenClaw. Pre-write guard before agents ship vague code. Zero deps. No model API.
[NeurIPS 2025] SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey 'LLM Agents: A Survey'.
Official repository for the paper "ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming"
A deterministic verification layer for AI systems. QWED verifies AI outputs using mathematics, symbolic reasoning, and formal methods (Z3, SMT, SymPy), creating an auditable trust boundary for agentic AI. Not generation. Verification.
Static analysis for AI agent configs, tool descriptions, and system prompts — catches vague tool descriptions, missing stop conditions, and schema gaps before they reach runtime. Zero-LLM, deterministic checks, built for CI.
Add a description, image, and links to the llm-safety topic page so that developers can more easily learn about it.
To associate your repository with the llm-safety topic, visit your repo's landing page and select "manage topics."