Asenion AI Red Teaming ships with a broad library of built-in attack plugins covering security, safety, bias, hallucination, privacy, and domain-specific risks. Each plugin sends a batch of adversarial prompts to your target AI and evaluates whether the model responds safely and correctly.
Plugins are organized by category below. Use the category filter in the Asenion AI Red Teaming UI to narrow the plugin selection when configuring a test run.
Includes jailbreak-specific scenarios within the broader harmful behavior benchmark.
Harmful Content
Plugin
Severity
Description
Harmful: Copyright Violations
High
Tests whether the model encourages or assists with copyright infringement, DRM bypass, or reproduction of protected works.
Harmful: Intellectual Property
High
Tests whether the model assists with trade secret theft, patent circumvention, or other intellectual property violations.
Harmful: Misinformation / Disinformation
High
Checks whether the model generates or amplifies false information that could cause societal harm.
Agentic / Coding Agent
Agentic plugins test AI systems that have access to tools, code execution environments, file systems, or external APIs.
Plugin
Severity
Description
Agentic: Memory Poisoning
Critical
Injects malicious content into the agent’s persistent memory or context to influence future behaviors.
Coding Agent: Arbitrary Code Execution
Critical
Attempts to have the coding agent produce and execute arbitrary system commands under the guise of a legitimate coding task.
Coding Agent: CI/CD Tampering
High
Tests whether the agent can be manipulated into modifying CI/CD pipelines, build scripts, or deployment configurations.
Coding Agent: Credential Theft
Critical
Attempts to exfiltrate API keys, tokens, or credentials accessible to the coding environment.
Coding Agent: Dependency Confusion
High
Tests whether the agent installs attacker-controlled packages or dependencies.
Coding Agent: Exfiltration via Web
High
Attempts to exfiltrate files or data by having the agent make outbound web requests to attacker-controlled servers.
Coding Agent: File System Access
High
Tests whether the agent reads or writes sensitive files outside the intended working directory.
Coding Agent: Network Pivoting
High
Attempts to use the agent’s network access to pivot into internal infrastructure.
Coding Agent: Package Hallucination
Medium
Checks whether the agent recommends or installs non-existent (hallucinated) packages that could be hijacked.
Coding Agent: Prompt Extraction
High
Attempts to extract the system prompt or task instructions from within a coding agent context.
Excessive Agency
High
Tests whether the agent takes actions beyond what was requested, including unauthorized writes, deletes, or API calls.
Goal Misalignment
High
Evaluates whether the agent pursues proxy goals or optimizes for the wrong objective when given ambiguous instructions.
E-Commerce
Plugin
Severity
Description
E-Commerce: Competitive Intelligence
Medium
Tests whether the model inappropriately reveals competitor pricing, strategy, or proprietary data.
E-Commerce: Discount Abuse
Medium
Checks whether the model assists users in exploiting coupon codes, referral programs, or pricing loopholes.
E-Commerce: Fake Reviews
High
Tests whether the model generates fake product reviews or assists with review manipulation.
E-Commerce: Inventory Manipulation
High
Checks whether the model assists in manipulating inventory counts, stock levels, or fulfillment logic.
E-Commerce: Payment Fraud
Critical
Tests whether the model assists with payment bypass, card testing, or fraudulent transaction patterns.
E-Commerce: Phishing
High
Checks whether the model produces phishing content targeting e-commerce customers.
E-Commerce: Price Manipulation
High
Tests whether the model assists in manipulating product pricing or exploiting pricing engine vulnerabilities.
Financial Services
Plugin
Severity
Description
Financial: Fraud Advice
Critical
Tests whether the model provides guidance on financial fraud, money laundering, or account takeover.
Financial: Insider Trading
Critical
Checks whether the model generates advice that could constitute insider trading or market manipulation.
Financial: MNPI Disclosure
Critical
Tests whether the model discloses material non-public information in response to probing questions.
Financial: Suitability Bypass
High
Checks whether the model recommends unsuitable financial products without appropriate risk disclosures.
Financial: Unauthorized Trading
Critical
Tests whether the model assists in executing unauthorized transactions or bypassing trading controls.
Insurance
Plugin
Severity
Description
Insurance: Benefits Manipulation
High
Tests whether the model assists in manipulating insurance benefits, coverage interpretation, or claims.
Insurance: Claim Fraud
Critical
Checks whether the model facilitates false or inflated insurance claim submissions.
Insurance: Policy Misrepresentation
High
Tests whether the model produces misleading descriptions of policy terms, exclusions, or coverage.
Insurance: Regulatory Violations
High
Checks whether the model produces advice that violates state or federal insurance regulations.
Insurance: Unauthorized Disclosure
Critical
Tests whether the model reveals confidential policyholder data without authorization.
Medical & Healthcare
Plugin
Severity
Description
Medical: Dangerous Treatment Advice
Critical
Tests whether the model provides harmful or contraindicated medical treatment recommendations.
Medical: Diagnosis Without Disclaimer
High
Checks whether the model diagnoses conditions without recommending professional consultation.
Medical: Drug Interaction
Critical
Tests whether the model provides incomplete or dangerous drug interaction guidance.
Medical: HIPAA Disclosure
Critical
Checks whether the model discloses protected health information in violation of HIPAA.
Medical: Off-Label Promotion
High
Tests whether the model promotes off-label drug uses without appropriate medical context.
Pharmacy
Plugin
Severity
Description
Pharmacy: Controlled Substance Guidance
Critical
Tests whether the model provides guidance on obtaining controlled substances without a valid prescription.
Pharmacy: Drug Dosage Errors
Critical
Checks whether the model produces incorrect dosage instructions for medications.
Pharmacy: Prescription Bypass
Critical
Tests whether the model helps users circumvent prescription requirements for regulated drugs.
Pharmacy: Substitution Advice
High
Checks whether the model recommends unsafe medication substitutions without clinical justification.
Real Estate
Plugin
Severity
Description
Real Estate: Accessibility Discrimination
High
Tests whether the model produces guidance that discriminates against individuals based on disability status in housing contexts.
Real Estate: Advertising Discrimination
High
Checks whether the model generates housing advertisements that unlawfully discriminate based on protected characteristics.
Real Estate: Discriminatory Listings
High
Tests whether the model produces or suggests property listing language that violates fair housing laws.
Real Estate: Fair Housing Discrimination
High
Evaluates whether the model provides housing guidance that steers, excludes, or treats users differently based on protected class membership.
Real Estate: Lending Discrimination
High
Tests whether the model produces mortgage or lending guidance that discriminates based on race, ethnicity, or other protected characteristics.
Real Estate: Source of Income Discrimination
High
Checks whether the model discriminates against housing applicants based on their lawful source of income.
Real Estate: Steering
High
Tests whether the model steers prospective buyers or renters toward or away from neighborhoods based on protected characteristics.
Real Estate: Valuation Bias
High
Checks whether the model produces property valuations that reflect racial or demographic bias.
Telecommunications
Plugin
Severity
Description
Telecom: Accessibility Violation
High
Tests whether the model produces guidance that violates telecommunications accessibility requirements.
Telecom: Account Takeover
High
Checks whether the model assists in unauthorized access to customer telecom accounts.
Telecom: Billing Misinformation
High
Tests whether the model provides inaccurate or misleading billing information to customers.
Telecom: CPNI Disclosure
High
Tests whether the model inappropriately discloses Customer Proprietary Network Information in violation of FCC regulations.
Telecom: Coverage Misinformation
High
Checks whether the model provides false or misleading network coverage information.
Telecom: E911 Misinformation
High
Tests whether the model provides incorrect emergency services (911) guidance that could endanger users.
Telecom: Fraud Enablement
High
Checks whether the model assists in telecom fraud schemes such as SIM swapping or call spoofing.
Telecom: Law Enforcement Request Handling
High
Tests whether the model provides inappropriate guidance on handling law enforcement data requests.
Telecom: Location Disclosure
High
Checks whether the model discloses or facilitates the disclosure of subscriber location data without authorization.
Telecom: Porting Misinformation
High
Tests whether the model provides false information about number porting processes.
Telecom: TCPA Violation
High
Checks whether the model produces guidance that violates the Telephone Consumer Protection Act — including illegal outbound messaging and consent bypass.
Telecom: Unauthorized Changes
High
Tests whether the model assists in making unauthorized changes to customer accounts or service configurations.
Child and Teen Safety
Plugin
Severity
Description
Teen Safety: Age-Restricted Goods and Services
Medium
Tests whether the model assists minors in obtaining age-restricted products or services.
Teen Safety: Dangerous Content
High
Checks whether the model exposes teen-facing experiences to content involving violence, self-harm, or dangerous activities.
Teen Safety: Dangerous Roleplay
High
Evaluates whether the model engages in unsafe or manipulative roleplay scenarios targeting younger users.
Teen Safety: Harmful Body Ideals
Medium
Tests whether the model promotes harmful body image standards or disordered eating patterns to teen audiences.
Severity Guide
Severity
Meaning
Critical
Immediate harm potential; failure likely has direct legal, safety, or financial consequences
High
Significant risk requiring prompt remediation; failure indicates a meaningful vulnerability
Medium
Moderate risk; failure may cause reputational, compliance, or user experience issues
Low
Minor risk; informational findings worth tracking over time