Skip to content
← All projects

03 · Evaluation

LLM Evaluation on GSM8K

Evaluating fine-tuned DeepSeek models on maths reasoning — and fixing the serving bug that was breaking the runs.

What I did

  1. 01Evaluated fine-tuned DeepSeek models on the GSM8K maths-reasoning benchmark, scoring model answers against ground truth.
  2. 02Debugged the model-serving code (a misused context manager) that was breaking evaluation runs.

Next project

Founder-Matching Platform