My Local LLM Scored 6/6. It Was Wrong Every Time.
Six months of trying to make a 1.2B model useful, and the measurement mistakes I made along the way.
Read articleWriting
Experiments, failures, and practical notes on local AI, agent systems, and trustworthy software.
Six months of trying to make a 1.2B model useful, and the measurement mistakes I made along the way.
Read article