← Back to all projectsAgentic AI & LLM Systems
Conversational LLM Fine-Tuning — SFT + DPO + LoRA/QLoRA
End-to-end, CPU-runnable fine-tuning pipeline that turns a small instruct model into a conversational signal classifier and compliant-response suggester: synthetic data generation, LoRA-based supervised fine-tuning, DPO preference optimization continuing the SFT adapter, an evaluation harness, and a merge/export step for serving, with a production fine-tuning and promotion playbook.
SFTDPOLoRAPEFTHuggingFace
View source on GitHub