PyMETA: Evaluating Student Code Diagnosis on and Beyond the First Execution Error

Publication
TAE (Trust-AI-Eval): Can We Trust AI Evaluation?
Create your slides in Markdown - click the Slides button to check out the example.

Supplementary notes can be added here, including code, math, and images.

Lingyu Gao
Lingyu Gao
AI Research Engineer

My work focuses on LLM-based systems for educational and multilingual applications.