RunLocalModel.com

Fine-Tune Qwen2.5 for Legal Work with LoRA and LLaMA-Factory

By the RunLocalModel editorial team · Published October 1, 2026 · ~11 minute read

If you only read one paragraph Fine-tuning takes a model that already speaks general language and teaches it a narrower job with examples. For Chinese legal questions, start from Qwen2.5-7B-Instruct, write a few hundred checked question-and-answer pairs, and train a small LoRA add-on in LLaMA-Factory. One 24GB GPU is enough. The model copies your examples. It does not grow a statute book.
Quick answers
What actually changes?
The habit of answering in the shape of your examples. Name the offense, walk the facts, cite the article. The reading ability was already in the base model.
Which file do I download?
The Instruct checkpoint, not Base. Instruct already knows how to take a question and answer it.
What data do I need?
An instruction, the case facts, and the answer you want copied. Mix in about 10–20% ordinary chat so normal conversation does not fall apart.
Full fine-tune or LoRA?
LoRA. You train a small patch and leave the original weights frozen. Rank 16 and alpha 32 is a normal first try on a 7B model.

This is the practical version of supervised fine-tuning. That page explains why showing the model an answer works. This one walks through one run: a legal assistant on Qwen2.5, trained with LoRA, on a single graphics card.

Start from a model that already follows instructions

The base model is the student. Extra legal files will not create reading ability or instruction-following that was never there. Qwen2.5-Instruct is a sensible start for Chinese legal work because the Instruct version already answers in a question-and-answer format. Your data budget goes to legal content, not to teaching the model what a conversation is.

Pick the size from the machine you actually have:

Model What you need When it makes sense
Qwen2.5-7B-Instruct One 24GB card, such as an RTX 3090 or 4090, or one A100 The first model to try. Long-context reading is already built in, up to 128k tokens, which matches contracts and case files. Training still uses a much shorter cutoff, covered below.
Qwen2.5-14B-Instruct One A100 or H100, or more than one 4090 Better at multi-step reasoning and at sticking to a complicated instruction. Useful once 7B answers are close but keep missing the logic.
Qwen2.5-32B or 72B Several A100s or H100s A production setting where the wording has to be tight and you can pay for the hardware.

Use the Instruct checkpoint unless you are doing continued pretraining on a very large pile of raw statute text, on the order of hundreds of gigabytes. The Base checkpoint only predicts the next word. It has not been taught to take an instruction and answer it. That extra teaching step is expensive, and most legal assistants do not need it. The difference between Base and Instruct is the same split described in pretraining and supervised fine-tuning.

The examples matter more than the training script

A fine-tune copies the pattern in your files. Vague, inconsistent, or wrong examples become vague, inconsistent, or wrong answers.

Three public sources are worth knowing. You can find them on Hugging Face or ModelScope:

Official text is the other half. Statutes from the National Laws and Regulations Database and judgments from China Judgments Online can be turned into the same instruction format. A statute or a judgment is not training data until you make it a question, the relevant facts, and a grounded answer.

One record, three fields

instruction is the task. input is the case. output is the answer you want the model to imitate, including the statute it should cite. If output is sloppy, the model will be sloppy in the same way.

[
  {
    "instruction": "Given the facts below, name the offense the defendant may have committed and explain why.",
    "input": "Facts: At night, Li entered a supermarket and took 3,000 yuan in cash plus goods worth 2,000 yuan.",
    "output": "Li's conduct constitutes theft. Under Article 264 of the Criminal Law of the PRC, taking public or private property with the intent of illegal possession, where the amount is relatively large, is theft."
  }
]

Mix in about 10–20% ordinary chat, such as ShareGPT-style or Alpaca-style conversations. Legal-only training pushes the model so hard toward that style that ordinary replies get worse. People call this catastrophic forgetting. The general examples keep normal conversation intact while the legal examples teach the new skill.

A few hundred clean, checked examples usually beat tens of thousands of scraped ones. Read a sample of the outputs yourself before you train.

What LoRA is doing

Full fine-tuning updates every weight. On a 7B model that is slow, it needs a lot of memory, and it can overwrite general ability you wanted to keep.

LoRA leaves the original weights frozen and trains a small pair of matrices beside them. At inference those matrices are added back in, so the model behaves as if it had been updated. You store and train only the small piece. The longer comparison with QLoRA and DoRA is on the supervised fine-tuning page.

The settings in the command below are a normal first try for a 7B Instruct model:

Run it in LLaMA-Factory

LLaMA-Factory wraps the training loop. You pass flags instead of writing the trainer yourself. Install it in its own environment:

git clone https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory

conda create -n llama_factory python=3.10 -y
conda activate llama_factory

pip install -e ".[torch,metrics]"
pip install flash-attn --no-build-isolation

FlashAttention-2 speeds up attention and uses less memory. If that install fails, training can still run without it. It will just use more memory.

Tell the loader where the fields are

Put legal_sft.json in data/, then add an entry to data/dataset_info.json. This mapping says which field is the task, which is the case, and which is the answer:

"legal_sft_data": {
  "file_name": "legal_sft.json",
  "columns": {
    "prompt": "instruction",
    "query": "input",
    "response": "output"
  }
}

The training command

On a single 24GB GPU:

CUDA_VISIBLE_DEVICES=0 llamafactory-cli train \
    --stage sft \
    --do_train true \
    --model_name_or_path Qwen/Qwen2.5-7B-Instruct \
    --dataset legal_sft_data \
    --template qwen \
    --finetuning_type lora \
    --lora_target all \
    --lora_rank 16 \
    --lora_alpha 32 \
    --lora_dropout 0.05 \
    --output_dir ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
    --overwrite_output_dir true \
    --cutoff_len 2048 \
    --preprocessing_num_workers 16 \
    --per_device_train_batch_size 2 \
    --gradient_accumulation_steps 4 \
    --learning_rate 1e-4 \
    --num_train_epochs 3.0 \
    --lr_scheduler_type cosine \
    --warmup_ratio 0.1 \
    --logging_steps 10 \
    --save_steps 100 \
    --bf16 true \
    --plot_loss true

The flags that decide whether the run fits on the card:

Watch the loss plot. Loss should fall and then flatten. If it falls to nearly zero in the first epoch, the model is memorizing. If it barely moves, the data or the learning rate is off.

Check the adapter before you merge it

The training run writes a LoRA adapter, not a full new model. Chat with that adapter first:

CUDA_VISIBLE_DEVICES=0 llamafactory-cli chat \
    --model_name_or_path Qwen/Qwen2.5-7B-Instruct \
    --adapter_name_or_path ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
    --template qwen

Ask questions that were not in the training file. Check three things: does it cite a real statute, does the reasoning follow the facts you gave, and can it still answer an ordinary non-legal question. A model that only sounds legal is a failed run.

When the answers are good enough, merge the adapter into the base weights. vLLM, Ollama, and similar servers can then load one normal model folder:

CUDA_VISIBLE_DEVICES=0 llamafactory-cli export \
    --model_name_or_path Qwen/Qwen2.5-7B-Instruct \
    --adapter_name_or_path ./saves/Qwen2.5-7B-Instruct/lora/legal_model \
    --template qwen \
    --export_dir ./exported_models/Qwen2.5-7B-Legal \
    --export_size 2 \
    --export_device cpu

export_size 2 splits the saved weights into 2GB pieces. export_device cpu does the merge on the CPU, so the GPU does not have to hold the full model during export.

How to read the result

Fine-tuning teaches style, format, and the patterns in your examples. It does not give the model a live database of statutes, and it does not stop it from stating a plausible but wrong article number. For anything you would rely on, pair the model with retrieval over the current statute text, and treat the generated answer as a draft.

If the first run is weak, change the data before you change the hyperparameters. Fix wrong outputs, drop duplicates, and add examples of the exact mistakes you are seeing. Rank, learning rate, and epoch count are the second lever.

Related guides on this site