Mitwa.ai
Fine-tuning Mistral-7B with DPO
With three college friends I built the backend of an AI wellbeing companion, then fine-tuned Mistral-7B when the base model wasn't good enough.
7BMistral, DPO fine-tuned
6k + 18kpreference pairs
4friends, one product
What happened
I built the backend, auth and OpenAI integration for an AI wellbeing companion. When the base model wasn't good enough, a friend and I fine-tuned Mistral-7B with DPO: for every prompt, a chosen answer and a rejected one, on datasets of 6k and 18k pairs.
Why it matters
It's where I learned that a model is only as good as the data and the measurement around it, which is the thread through everything I've built since.
Stack
PythonHugging Face TRLMistral-7BOpenAI API