From Evaluation to Guardrails: What We Brought to ACM FAccT 2026
# Summary
Researchers presented a tutorial at ACM FAccT 2026 arguing that guardrails—mechanisms filtering LLM inputs and outputs—deserve the same rigorous evaluation as the models themselves. They demonstrated this by evaluating refugee and asylum-focused scenarios across five languages with native speakers, then converting identified failures into concrete guardrail policies in English and Farsi. The work revealed that text-only guardrails have fundamental limitations in verifying factuality and context-specific accuracy, suggesting that effective guardrails require LLM-enabled verification capabilities.
Read Full Article →