SEQUOR is a benchmark for evaluating LLMs on multi-turn realistic constraint following. It includes five test sets — single, tuples, replace, add, and everything — built from a fixed collection of user-turn sequences that differ in how constraints are introduced and modified across conversation turns. The repository includes the full pipeline for constraint collection and filtering, testset generation, model response generation, evaluation with LLM-as-a-judge, and result analysis.
📄 Paper: https://openreview.net/forum?id=93Xw0UkZhC
SEQUOR was accepted at COLM 2026 🎉🎉
data/: Constraints, tasks, and test sets that constitute SEQUOR. Also contains the gold responses used to validate the LLM-as-a-judge evaluation.pipeline/: Constraint collection and filtering pipeline (extraction, deduplication, satisfiability, triviality, subjectivity, tuple creation) and evaluation with LLM-as-a-judge. Seepipeline/README.md.multi_if/: Testset generation, model response generation, scoring, and plotting. Seemulti_if/README.md.
If you use SEQUOR or this repository in your work, please cite our paper:
@inproceedings{canaverde2026sequor,
title={{SEQUOR}: A Multi-Turn Benchmark for Realistic Constraint Following},
author={Beatriz Canaverde and Duarte Miguel Alves and Jos{\'e} Pombal and Giuseppe Attanasio and Andre Martins},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=93Xw0UkZhC}
}