03 · Evaluation
LLM Evaluation on GSM8K
Evaluating fine-tuned DeepSeek models on maths reasoning — and fixing the serving bug that was breaking the runs.
What I did
- 01Evaluated fine-tuned DeepSeek models on the GSM8K maths-reasoning benchmark, scoring model answers against ground truth.
- 02Debugged the model-serving code (a misused context manager) that was breaking evaluation runs.
Next project
Founder-Matching Platform