Counterfactual Training

Written by

Counterfactual training is a machine learning technique that improves models by training them on “what-if” scenarios—hypothetical alternative inputs and their desired outputs—to enhance robustness, fairness, and explainability.

It is often used to reduce bias in AI systems, enhance safety in autonomous vehicles, and improve the reliability of large language models (LLMs) by ensuring faithful reasoning.