This page provides the security targets of negative impacts “Improper manipulation of consumer decision-making by AI” in the external influence aspect in the AI Security Map, as well as the attacks and factors that cause them, and the corresponding defense methods and countermeasures.
Security target
- Consumer
Attack or cause
- Integrity violation
- Degradation of output fairness
- Degradation of controllability
- Reliability violation
Defensive method or countermeasure
- Human in the loop
- AI alignment
- XAI (Explainable AI)
- Uncertainty quantification
References
Human in the loop
AI alignment
- Training language models to follow instructions with human feedback, 2022
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, 2022
- Constitutional AI: Harmlessness from AI Feedback, 2022
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model, 2023
- A General Theoretical Paradigm to Understand Learning from Human Preferences, 2023
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears, 2023
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, 2023
- Self-Rewarding Language Models, 2024
- KTO: Model Alignment as Prospect Theoretic Optimization, 2024
- SimPO: Simple Preference Optimization with a Reference-Free Reward, 2024
XAI (Explainable AI)
- Visualizing and Understanding Convolutional Networks, 2014
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps, 2014
- Understanding Deep Image Representations by Inverting Them, 2014
- “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, 2016
- A Unified Approach to Interpreting Model Predictions, 2017
- Learning Important Features Through Propagating Activation Differences, 2017
- Understanding Black-box Predictions via Influence Functions, 2017
- Interpretable Explanations of Black Boxes by Meaningful Perturbation, 2017
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV), 2018
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization, 2019