← Rishi Sangare
Mitwa.ai

Fine-tuning Mistral-7B with DPO

With three college friends I built the backend of an AI wellbeing companion, then fine-tuned Mistral-7B when the base model wasn't good enough.

Backend and ML engineerDec 2023 – Apr 2024
7BMistral, DPO fine-tuned
6k + 18kpreference pairs
4friends, one product

What happened

I built the backend, auth and OpenAI integration for an AI wellbeing companion. When the base model wasn't good enough, a friend and I fine-tuned Mistral-7B with DPO: for every prompt, a chosen answer and a rejected one, on datasets of 6k and 18k pairs.

Why it matters

It's where I learned that a model is only as good as the data and the measurement around it, which is the thread through everything I've built since.

Stack

PythonHugging Face TRLMistral-7BOpenAI API
nextThe Studio: an AI content engine →