← Back to all projectsAgentic AI & LLM Systems

Conversational LLM Fine-Tuning — SFT + DPO + LoRA/QLoRA

End-to-end, CPU-runnable fine-tuning pipeline that turns a small instruct model into a conversational signal classifier and compliant-response suggester: synthetic data generation, LoRA-based supervised fine-tuning, DPO preference optimization continuing the SFT adapter, an evaluation harness, and a merge/export step for serving, with a production fine-tuning and promotion playbook.

SFTDPOLoRAPEFTHuggingFace
View source on GitHub