Fine-Tune Qwen2.5 for Legal Work with LoRA and LLaMA-Factory
- What actually changes?
- The habit of answering in the shape of your examples. Name the offense, walk the facts, cite the article. The reading ability was already in the base model.
- Which file do I download?
- The Instruct checkpoint, not Base. Instruct already knows how to take a question and answer it.
- What data do I need?
- An instruction, the case facts, and the answer you want copied. Mix in about 10–20% ordinary chat so normal conversation does not fall apart.
- Full fine-tune or LoRA?
- LoRA. You train a small patch and leave the original weights frozen. Rank 16 and alpha 32 is a normal first try on a 7B model.
This is the practical version of supervised fine-tuning. That page explains why showing the model an answer works. This one walks through one run: a legal assistant on Qwen2.5, trained with LoRA, on a single graphics card.
Start from a model that already follows instructions
The base model is the student. Extra legal files will not create reading ability or instruction-following that was never there. Qwen2.5-Instruct is a sensible start for Chinese legal work because the Instruct version already answers in a question-and-answer format. Your data budget goes to legal content, not to teaching the model what a conversation is.
Pick the size from the machine you actually have:
| Model | What you need | When it makes sense |
|---|---|---|
| Qwen2.5-7B-Instruct | One 24GB card, such as an RTX 3090 or 4090, or one A100 | The first model to try. Long-context reading is already built in, up to 128k tokens, which matches contracts and case files. Training still uses a much shorter cutoff, covered below. |
| Qwen2.5-14B-Instruct | One A100 or H100, or more than one 4090 | Better at multi-step reasoning and at sticking to a complicated instruction. Useful once 7B answers are close but keep missing the logic. |
| Qwen2.5-32B or 72B | Several A100s or H100s | A production setting where the wording has to be tight and you can pay for the hardware. |
Use the Instruct checkpoint unless you are doing continued pretraining on a very large pile of raw statute text, on the order of hundreds of gigabytes. The Base checkpoint only predicts the next word. It has not been taught to take an instruction and answer it. That extra teaching step is expensive, and most legal assistants do not need it. The difference between Base and Instruct is the same split described in pretraining and supervised fine-tuning.
The examples matter more than the training script
A fine-tune copies the pattern in your files. Vague, inconsistent, or wrong examples become vague, inconsistent, or wrong answers.
Three public sources are worth knowing. You can find them on Hugging Face or ModelScope:
- DISC-LawLLM (Fudan): Chinese legal instructions covering questions, judicial-exam items, and document drafting.
- CAIL: the China AI and Law competition sets. Structured tasks such as charge prediction, statute recommendation, and sentence length.
- ChatLaw and LaWGPT: instruction sets built around legal questions and answers.
Official text is the other half. Statutes from the National Laws and Regulations Database and judgments from China Judgments Online can be turned into the same instruction format. A statute or a judgment is not training data until you make it a question, the relevant facts, and a grounded answer.
One record, three fields
instruction is the task. input is the case. output is the answer you want the model to imitate, including the statute it should cite. If output is sloppy, the model will be sloppy in the same way.
[
{
"instruction": "Given the facts below, name the offense the defendant may have committed and explain why.",
"input": "Facts: At night, Li entered a supermarket and took 3,000 yuan in cash plus goods worth 2,000 yuan.",
"output": "Li's conduct constitutes theft. Under Article 264 of the Criminal Law of the PRC, taking public or private property with the intent of illegal possession, where the amount is relatively large, is theft."
}
]
Mix in about 10–20% ordinary chat, such as ShareGPT-style or Alpaca-style conversations. Legal-only training pushes the model so hard toward that style that ordinary replies get worse. People call this catastrophic forgetting. The general examples keep normal conversation intact while the legal examples teach the new skill.
A few hundred clean, checked examples usually beat tens of thousands of scraped ones. Read a sample of the outputs yourself before you train.
What LoRA is doing
Full fine-tuning updates every weight. On a 7B model that is slow, it needs a lot of memory, and it can overwrite general ability you wanted to keep.
LoRA leaves the original weights frozen and trains a small pair of matrices beside them. At inference those matrices are added back in, so the model behaves as if it had been updated. You store and train only the small piece. The longer comparison with QLoRA and DoRA is on the supervised fine-tuning page.
The settings in the command below are a normal first try for a 7B Instruct model:
lora_rank 16sets how wide that small update is. A larger rank can fit more new detail and costs more memory.lora_alpha 32scales how strongly the update affects the output. A common rule of thumb is alpha around twice the rank.lora_dropout 0.05randomly drops a bit of the adapter during training so it does not memorize the set.lora_target allattaches the adapters to the linear layers, so the patch can see the whole network.
Run it in LLaMA-Factory
LLaMA-Factory wraps the training loop. You pass flags instead of writing the trainer yourself. Install it in its own environment:
git clone https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory
conda create -n llama_factory python=3.10 -y
conda activate llama_factory
pip install -e ".[torch,metrics]"
pip install flash-attn --no-build-isolation
FlashAttention-2 speeds up attention and uses less memory. If that install fails, training can still run without it. It will just use more memory.
Tell the loader where the fields are
Put legal_sft.json in data/, then add an entry to data/dataset_info.json. This mapping says which field is the task, which is the case, and which is the answer:
"legal_sft_data": {
"file_name": "legal_sft.json",
"columns": {
"prompt": "instruction",
"query": "input",
"response": "output"
}
}
The training command
On a single 24GB GPU:
CUDA_VISIBLE_DEVICES=0 llamafactory-cli train \
--stage sft \
--do_train true \
--model_name_or_path Qwen/Qwen2.5-7B-Instruct \
--dataset legal_sft_data \
--template qwen \
--finetuning_type lora \
--lora_target all \
--lora_rank 16 \
--lora_alpha 32 \
--lora_dropout 0.05 \
--output_dir ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
--overwrite_output_dir true \
--cutoff_len 2048 \
--preprocessing_num_workers 16 \
--per_device_train_batch_size 2 \
--gradient_accumulation_steps 4 \
--learning_rate 1e-4 \
--num_train_epochs 3.0 \
--lr_scheduler_type cosine \
--warmup_ratio 0.1 \
--logging_steps 10 \
--save_steps 100 \
--bf16 true \
--plot_loss true
The flags that decide whether the run fits on the card:
--stage sftis supervised fine-tuning. The model learns to produceoutputgiveninstructionandinput.--template qwenwraps each example in Qwen’s chat markers. The wrong template trains the model on text it will never see when you chat with it.--cutoff_len 2048clips each example to 2048 tokens. Qwen2.5 can read far more than that, but training at 8k or 32k on a 24GB card runs out of memory. Raise this only after a short run succeeds, and only if your examples are actually that long.- Batch size 2 with 4 accumulation steps is an effective batch of 8. If you run out of memory, drop the batch size to 1 and raise the accumulation steps so the effective batch stays in the same range.
- Learning rate
1e-4is a normal LoRA rate. It is higher than a full fine-tune rate because you are updating a small adapter. - Three epochs, a cosine schedule, and a 10% warmup are a standard first recipe. Warmup raises the learning rate gradually so the first updates are not huge.
--bf16 trueis the steadier choice on a 3090 or 4090. Bfloat16 keeps a wide exponent range, so the training is less likely to blow up than float16. Use one precision flag, not both.
Watch the loss plot. Loss should fall and then flatten. If it falls to nearly zero in the first epoch, the model is memorizing. If it barely moves, the data or the learning rate is off.
Check the adapter before you merge it
The training run writes a LoRA adapter, not a full new model. Chat with that adapter first:
CUDA_VISIBLE_DEVICES=0 llamafactory-cli chat \
--model_name_or_path Qwen/Qwen2.5-7B-Instruct \
--adapter_name_or_path ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
--template qwen
Ask questions that were not in the training file. Check three things: does it cite a real statute, does the reasoning follow the facts you gave, and can it still answer an ordinary non-legal question. A model that only sounds legal is a failed run.
When the answers are good enough, merge the adapter into the base weights. vLLM, Ollama, and similar servers can then load one normal model folder:
CUDA_VISIBLE_DEVICES=0 llamafactory-cli export \
--model_name_or_path Qwen/Qwen2.5-7B-Instruct \
--adapter_name_or_path ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
--template qwen \
--export_dir ./exported_models/Qwen2.5-7B-Legal \
--export_size 2 \
--export_device cpu
export_size 2 splits the saved weights into 2GB pieces. export_device cpu does the merge on the CPU, so the GPU does not have to hold the full model during export.
How to read the result
Fine-tuning teaches style, format, and the patterns in your examples. It does not give the model a live database of statutes, and it does not stop it from stating a plausible but wrong article number. For anything you would rely on, pair the model with retrieval over the current statute text, and treat the generated answer as a draft.
If the first run is weak, change the data before you change the hyperparameters. Fix wrong outputs, drop duplicates, and add examples of the exact mistakes you are seeing. Rank, learning rate, and epoch count are the second lever.
Related guides on this site
- Supervised Fine-Tuning Explained — why example answers work, plus LoRA, QLoRA, and DoRA
- How LLMs Are Trained — where fine-tuning sits in the whole process
- LLM Pretraining and Continued Pretraining — when you would use a Base checkpoint
- Choosing Quantization — how to shrink the merged model so it fits for chatting
- How to Run Local AI