Hi all, I’m a university researcher building an on-device assistive system for blind pedestrians (Raspberry Pi 5 + Hailo-10H): the user speaks a destination, and a VLM reads indoor wayfinding signs and returns only the direction for that destination as one label out of 14 fixed classes. From this thread ( Is there a documented flow for compiling a custom/fine-tuned LLM to the Hailo-10H LLM pipeline? ) I understand that compiling your own LLM is not possible yet, but LoRA on supported models is. Before I invest in fine-tuning, I’d like to confirm how far this applies to the VLMs in the GenAI Model Zoo: 1. Is LoRA supported for the language model part of Qwen2-VL-2B-Instruct and/or Qwen3-VL-2B-Instruct? If so, is a LoRA-ready HAR provided for these VLMs (as for Qwen2-1.5B), and which DFC / HailoRT version is required? (I have DFC 5.4.0 on x86.) 2. If supported: must the vision encoder stay exactly as shipped, with only the LLM adapter changing? Are there constraints on rank, target modules, number of adapters, or on the base checkpoint/quantization the adapter must be trained against? 3. Can the hailo_platform.genai VLM class load a HEF compiled with a LoRA adapter, or is the low-level HailoRT API needed (cf. Qwen compiled LoRA hef successful but not possible to create LLM object via hailo_platform.genai )? 4. Does the GenAI VLM API expose next-token logits/probabilities, or allow restricting generation to a list of allowed tokens (constrained decoding)? We need the probability of each of our 14 single-token labels to decide when to abstain. 5. If VLM LoRA isn’t supported yet, is it on the roadmap? And separately, can a fine-tuned vision encoder alone be compiled with the DFC (the Model Zoo lists standalone encoders such as Qwen2-VL-2B-vision-336x336)? Thanks in advance!
1 Like