The Model Agreed, But Didn’t Learn: Diagnosing Surface Compliance in Large Language Models
Published in Findings of ACL 2026, 2026
Co-first authors.
This paper introduces SA-MCQ, a discriminative diagnostic framework for revealing whether edited knowledge is truly internalized by large language models.
Recommended citation: Gu, X.*, Huang, Z.*, Hong, W., Xie, J., Lou, R., & Zhang, K. (2026). The Model Agreed, But Didn’t Learn: Diagnosing Surface Compliance in Large Language Models. Findings of ACL 2026.
Download Paper
