Hi,
I’ve been trying to get a small LoRA adapter running on a Raspberry Pi 5 with the AI HAT+ 2, following DFC_7_LoRA_Tutorial pretty much to the letter. The good part is the whole thing works end to end: the HEF compiles, loads, and generates text. The bad part is that on the device the adapter makes the model worse than no adapter at all, and I’ve run out of things to check on my side.
Setup
- Pi 5 8GB + AI HAT+ 2, Raspberry Pi OS (trixie), HailoRT and firmware 5.1.1 (newest in the RPi apt repo)
- DFC 5.3.0, qwen2_1.5b_instruct.q.har and both .alls files from dev-public v5.3.0
- LoRA r=32, alpha=64, gate/up/down only, trained with PEFT/TRL on about 2000 examples, calibset 64 like in the tutorial
- compiled on Ubuntu 24.04 (WSL2), roughly 6 hours total
- running it through hailo-ollama (model zoo 5.1.1), I registered the HEF under a new manifest copied from qwen2:1.5b. I also tried hailo_platform.genai.LLM directly with lora_name set to my adapter name, output is identical
Task and results
The task is simple: write one sentence about home network status and copy the area names from the prompt exactly. I score 120 held-out prompts with a rule-based checker, greedy decoding, max 60 tokens.
GPU (fp) Hailo-10H
Qwen2-1.5B-Instruct, no adapter 74.2 % 55.8 % (qwen2:1.5b from the model zoo)
Qwen2-1.5B-Instruct + my LoRA 100.0 % 39.2 % (my HEF)
So on GPU the adapter fixes the task completely, on the chip it drops below the stock model. Most of the device errors are small corruptions of names copied from the prompt: “ADS-B” → “AD-B”, “backups” → “back-ups”, “Discord bots” → “Discord bot”, missing spaces like “andADS-B” or “fine!honeypot”. Output is deterministic and stops where it should, so it’s not sampling or stop tokens. Setting repetition_penalty to 1.0 in the manifest changed nothing either.
What I checked
My first guess was plain 4-bit quantization, so I tried to reproduce the quantized model in PyTorch from the HAR itself:
- weights: dequantized the w4 group-128 kernels (group scales from output_stage slopes_m/slopes_e), undid the QuaRot rotation and the folded RMSNorm gammas. All 7 linear types come out with ~10.4% relative error vs the fp weights, which looks right for 4 bit
- 8-bit inputs of all linear layers using qp_in, 16-bit input for down_proj
- 8-bit V and 8-bit exp/softmax probabilities, following the softmax layers in the HN graph
With that, the base model scores 56.7%, which is basically what the chip does (55.8%), so I think the simulation is close enough. (The full simulation is slow, so these are on a 30 and 60 prompt subset, not all 120.) The same simulation with my adapter gives ~83%. That’s with the base activation ranges and also with the recalibrated ranges from the adapter’s network group in my optimized HAR. So on paper the adapter should survive quantization fine, but on the device it loses another ~45 points.
A few other things I verified:
- load_lora_weights does what I’d expect: the norm gamma is folded into lora_down and alpha/r = 2 into lora_up, the fpo weights match my adapter to ~1e-8
- adapter strength isn’t off. I compared the device outputs to the simulation at 0x-4x LoRA scale and 1x matches best
- the adapter style shows up in the whole generated sentence, so both prefill and tbt seem to use it
Questions
- My HEF is from DFC 5.3.0 but the Pi only has HailoRT 5.1.1 in the apt repo. I checked whether that matters: the official Qwen2-1.5B-Instruct.hef from dev-public v5.3.0 scores 55.0% on my 5.1.1 runtime, same as the 5.1.1 HEF (55.8%). So the version mismatch doesn’t seem to hurt the base model. Is there anything LoRA-specific in the runtime that changed between 5.1 and 5.3, though?
- Is there a supported way to run the LoRA model in quantized emulation (SDK_QUANTIZED) for the LLM flow? That would tell me if the drop happens at optimization/compile time or at runtime.
- Has anyone measured task accuracy of a LoRA HEF on the device vs GPU? The tutorial only reports GPU numbers, and I couldn’t find anything here on the forum.
I can share the adapter, the scripts and the test set if someone wants to reproduce it.
Thanks!