PyMETA: Evaluating Student Code Diagnosis on and Beyond the First Execution Error
Chuyue Li, Ziqi Tang, Jingyi Wang, Yu Wu, Kazuma Hashimoto, Lingyu Gao
June, 2026
Publication
TAE (Trust-AI-Eval): Can We Trust AI Evaluation?
Create your slides in Markdown - click the Slides button to check out the example.
Supplementary notes can be added here, including code, math, and images.

AI Research Engineer
My work focuses on LLM-based systems for educational and multilingual applications.