Hailo-10H / HailoRT 5.4.0: ConfiguredInferModel::run_async_for_duration() always fails with HAILO_RPC_FAILED (77) — host seems to parse the reply from the header, not the body

On Hailo-10H, run_async_for_duration() fails on every call with status 77, after waiting the full requested duration. From reading the v5.4.0 source, I think the host-side reply parsing reads from the wrong offset. Details, a minimal reproduction and a suggested one-line fix are below.

Disclosure: This was found by my AI coding agent (Claude, running in Claude Code) while it was benchmarking my Raspberry Pi 5 + Hailo-10H setup. The agent wrote the test program, read the HailoRT v5.4.0 source, did the offline protobuf check, and ran the device tests on my Pi with my approval. I reviewed this write-up before posting. In the text below, “I” covers that work.

Environment

  • Hardware: Raspberry Pi 5 (16 GB) + Raspberry Pi AI HAT+ 2 (Hailo-10H)

  • OS: Raspberry Pi OS (Debian 13 “trixie”), 64-bit

  • Kernel: 6.18.50+rpt-rpi-2712 (16K pages)

  • HailoRT: 5.4.0 (libhailort.so.5.4.0)

  • PCIe driver: hailo1x_pci 5.4.0

  • Device firmware: 5.4.0 (release, app)

  • HEFs: scrfd_2.5g and scrfd_10g from Model Zoo v5.4.0; scrfd_10g and yolov8m_pose that were already installed on the system (compiler version 5.1.0 per hailortcli parse-hef)

  • Test program: g++ 14.2.0, -std=c++17

What I did

Goal: measure on-chip throughput without per-frame host-device transfers.

  1. VDevice::create()

  2. create_infer_model(hef) → configure() → create_bindings()

  3. Page-aligned (mmap) input and output buffers, attached with set_buffer()

  4. ConfiguredInferModel::run_async_for_duration(bindings, 3000, 0, callback)

  5. job->wait(duration + 15 s)

Minimal reproduction

chip_fps.cpp:

// chip_fps: measure on-chip throughput with ConfiguredInferModel::run_async_for_duration.

// On Hailo-10H the server loops on a single frame inside the chip; the host sends the

// input once at the start and receives the result once at the end (no per-frame PCIe).

#include "hailo/hailort.hpp"

#include <sys/mman.h>

#include <atomic>

#include <chrono>

#include <cstdlib>

#include <cstring>

#include <iostream>

#include <string>

using namespace hailort;

static void *alloc_page_aligned(size_t size)

{

void *p = mmap(nullptr, size, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);

return (MAP_FAILED == p) ? nullptr : p;

}

int main(int argc, char **argv)

{

if (argc < 2) {

std::cerr << "usage: chip_fps <hef> [duration_ms] [sleep_between_frames_ms]" << std::endl;

return 2;

}

const std::string hef_path = argv[1];

const uint32_t duration_ms = (argc > 2) ? static_cast<uint32_t>(std::atoi(argv[2])) : 3000;

const uint32_t sleep_ms = (argc > 3) ? static_cast<uint32_t>(std::atoi(argv[3])) : 0;

auto vdevice = VDevice::create();

if (!vdevice) { std::cerr << "VDevice::create failed: " << vdevice.status() << std::endl; return 1; }

auto infer_model_exp = vdevice.value()->create_infer_model(hef_path);

if (!infer_model_exp) { std::cerr << "create_infer_model failed: " << infer_model_exp.status() << std::endl; return 1; }

auto infer_model = infer_model_exp.release();

auto configured_exp = infer_model->configure();

if (!configured_exp) { std::cerr << "configure failed: " << configured_exp.status() << std::endl; return 1; }

auto configured = configured_exp.release();

auto bindings_exp = configured.create_bindings();

if (!bindings_exp) { std::cerr << "create_bindings failed: " << bindings_exp.status() << std::endl; return 1; }

auto bindings = bindings_exp.release();

size_t in_bytes = 0;

size_t out_bytes = 0;

for (const auto &name : infer_model->get_input_names()) {

const size_t size = infer_model->input(name)->get_frame_size();

void *p = alloc_page_aligned(size);

if (nullptr == p) { std::cerr << "alloc failed" << std::endl; return 1; }

std::memset(p, 0x40, size);

auto status = bindings.input(name)->set_buffer(MemoryView(p, size));

if (HAILO_SUCCESS != status) { std::cerr << "set input buffer failed: " << status << std::endl; return 1; }

in_bytes += size;

}

for (const auto &name : infer_model->get_output_names()) {

const size_t size = infer_model->output(name)->get_frame_size();

void *p = alloc_page_aligned(size);

if (nullptr == p) { std::cerr << "alloc failed" << std::endl; return 1; }

auto status = bindings.output(name)->set_buffer(MemoryView(p, size));

if (HAILO_SUCCESS != status) { std::cerr << "set output buffer failed: " << status << std::endl; return 1; }

out_bytes += size;

}

std::atomic<uint32_t> fps{0};

std::atomic<int> callback_status{-1};

const auto t0 = std::chrono::steady_clock::now();

auto job = configured.run_async_for_duration(bindings, duration_ms, sleep_ms,

[&fps, &callback_status](const AsyncInferCompletionInfo &info, uint32_t measured_fps) {

callback_status = static_cast<int>(info.status);

fps = measured_fps;

});

if (!job) { std::cerr << "run_async_for_duration failed: " << job.status() << std::endl; return 1; }

const auto wait_status = job->wait(std::chrono::milliseconds(duration_ms + sleep_ms + 15000));

const auto t1 = std::chrono::steady_clock::now();

const auto wall_ms = std::chrono::duration_cast<std::chrono::milliseconds>(t1 - t0).count();

std::cout << "sleep_ms=" << sleep_ms << " chip_loop_fps=" << fps.load() << " callback_status=" << callback_status.load()

<< " wait_status=" << wait_status << " wall_ms=" << wall_ms

<< " input_bytes=" << in_bytes << " output_bytes=" << out_bytes << std::endl;

return (HAILO_SUCCESS == wait_status && 0 == callback_status.load()) ? 0 : 1;

}

Build and run:

g++ -O2 -std=c++17 chip_fps.cpp -o chip_fps -lhailort -lpthread

./chip_fps scrfd_2.5g.hef 3000 # sleep_between_frames_ms = 0 -> fails with 77

./chip_fps scrfd_2.5g.hef 3000 5000 # fps rounds down to 0 -> succeeds (control)

Expected

Callback status HAILO_SUCCESS, job->wait() returns HAILO_SUCCESS, and fps > 0.

Actual

  • 9 of 9 runs failed, across 4 HEFs: Model Zoo v5.4.0 scrfd_2.5g (3 runs) and scrfd_10g (3 runs), pre-installed scrfd_10g (2 runs) and yolov8m_pose (1 run).

  • Host output on every run:

    [HailoRT] [error] CHECK failed - Failed to de-serialize 'RunAsyncForDuration'
    
  • Callback status: 77 (HAILO_RPC_FAILED). job->wait() also returns 77.

  • Wall time from the call to wait() returning: 3,029–3,042 ms. It fails only after the requested 3,000 ms.

  • The host does not print “Calling run async for duration has failed”, which is what it logs when the server replies with an error status.

Control experiment: it succeeds only when fps is 0

If the host parses the reply from the header (see the analysis below), the call should succeed only when the reply body is empty. Protobuf omits fps when it is 0, so header.size becomes 0. I forced this with sleep_between_frames_ms (same program, scrfd_2.5g, duration_ms = 3000):

sleep_between_frames_ms frames in 3 s → fps sent by device reply body prediction result
0 ~720 → ~240 3 bytes 77 77 (wall 3,031 ms)
1000 3 → 1 2 bytes (08 01) 77 77 (wall 3,006 ms)
5000 1 → 0 (0.33 truncated) 0 bytes 0 0 (wall 5,005 ms), 2 of 2 runs

So the call fails whenever the device reports a non-zero fps and succeeds only when it reports 0. This matches the header-offset explanation and rules out a general RPC or device-side failure.

What works

Same HEFs, same board: hailortcli run2 (full_async) runs normally, e.g. scrfd_2.5g at about 240 fps.

Device log

hailortcli logs runtime for the same time window has no lines containing ForDuration or serialize. The only HailoRT-Server warnings in that window are descriptor/CCB warnings. I can post them if useful.

Source analysis (tag v5.4.0, commit f51959034a)

Paths are relative to hailort/ in the repository. This is my reading of the source, an offline protobuf check, and the on-device control experiment above. I have not rebuilt libhailort with the change, so the fix itself is untested.

Reply path on Hailo-10H (hRPC)

  • Host: ConfiguredInferModel::run_async_for_duration (libhailort/src/net_flow/pipeline/infer_model.cpp:707-711) forwards to ConfiguredInferModelHrpcClient::run_async_for_duration (libhailort/src/net_flow/pipeline/configured_infer_model_hrpc_client.cpp:614-672).

  • Device: ConfiguredInferModelRunAsyncForDurationHandler (hailort_server/hailort_server.cpp:746-840) calls ConfiguredInferModelImpl::run_async_for_duration (infer_model.cpp:915-959). That loops run_async on one frame for duration_ms and then reports fps. The handler serializes uint32 fps = 1 (hrpc_protocol/rpc.proto:194-197, hrpc_protocol/serializer.cpp:764-774). ResponseWriter::write sends header + body as one message, then the output buffers as separate transfers (hrpc/response_writer.hpp:29-73).

  • Host: Client::message_loop (hrpc/client.cpp:84-112) calls RpcConnection::parse_message (hrpc/rpc_connection.cpp:73-77). It returns body = buffer->data() + sizeof(rpc_message_header_t) with length header.size. The callback receives rpc_message_t{buffer, header, body} (hrpc/rpc_connection.hpp:39-44, hrpc/client.cpp:107).

Likely cause — confidence: high

configured_infer_model_hrpc_client.cpp:648 (link):

auto expected_fps = RunAsyncForDurationSerializer::deserialize_reply(MemoryView(reply.buffer->data(), reply.header.size));

reply.buffer->data() is offset 0 of the message, which is the 20-byte packed header (magic, size, message_id, action_id, status; hrpc/rpc_connection.hpp:28-37). The body starts at offset 20. So the parser gets the first header.size bytes of the header.

RPC_MESSAGE_MAGIC is 0x8A554432 (rpc_connection.hpp:22), little-endian bytes 32 44 55 8A. Protobuf reads 0x32 as field 6, wire type 2 (length-delimited), and 0x44 as a length of 68. The reply body is only 2–6 bytes (fps = 240 → 08 F0 01, 3 bytes), so ParseFromArray fails. deserialize_reply then returns HAILO_RPC_FAILED with the exact message above (serializer.cpp:775-780).

Other reply parses use the body, for example:

  • configured_infer_model_hrpc_client.cpp:410: MemoryView(result.body.data(), result.header.size)

  • infer_model_hrpc_client.cpp:112: same pattern

As far as I can see, line 648 is the only reply parse in the tree that uses reply.buffer->data().

This matches what I observed:

  • It fails every time fps > 0. Only fps == 0 (empty body, header.size == 0) parses — confirmed on the device by the control experiment.

  • Wall time ≈ requested duration: the device ran the whole loop and replied with success.

  • Nothing in the device log, and no “Calling run async for duration has failed” on the host: the server parsed the request and serialized the reply without error. The failure happens on the host after the reply arrives.

  • run_async() / run2 are not affected: the normal RunAsync handler sends an empty body (hailort_server.cpp:728) and its host callback does not parse one.

Offline check: with Python protobuf, google.protobuf.UInt32Value has the same wire format as uint32 fps = 1. I built header + body and parsed the first header.size bytes from offset 0. It fails (“Wire format was corrupt”) for fps = 1, 50, 127, 128, 240, 300, 16384, 2^21 and 2^28. Parsing from offset 20 gives the correct value. fps = 0 parses either way.

Ruled out — confidence: high

  • The device failing to parse the request. The same message text exists in deserialize_request (serializer.cpp:747). But when request parsing fails, the server replies right away with an error header (hrpc/server.cpp:150-159). Wall time would then be near 0, and the host would log “Calling run async for duration has failed” (configured_infer_model_hrpc_client.cpp:655). Neither happened.

  • Output data mixed into the reply body. header.size counts only the protobuf body (response_writer.hpp:38). Outputs follow as separate transfers (response_writer.hpp:64-70), and the host reads them into the user buffers (client.cpp:102-105). This is the same contract as the working RunAsync.

Suggested fix (host libhailort only; firmware and HailoRT-Server unchanged)

--- a/hailort/libhailort/src/net_flow/pipeline/configured_infer_model_hrpc_client.cpp

+++ b/hailort/libhailort/src/net_flow/pipeline/configured_infer_model_hrpc_client.cpp

@@ -648,1 +648,1 @@

- auto expected_fps = RunAsyncForDurationSerializer::deserialize_reply(MemoryView(reply.buffer->data(), reply.header.size));

+ auto expected_fps = RunAsyncForDurationSerializer::deserialize_reply(MemoryView(reply.body.data(), reply.header.size));

reply.body points into reply.buffer, which the callback holds, so there is no lifetime issue.

Still present in newer code

The same line is in v5.1.1 (line 656, where the function first appears), v5.2.0 (651), v5.3.0 (648), v5.4.0 (648), current master (648), and commit d50fdb9c5d from the closed “v5.5.0” PR (hailo-ai/hailort#45). If 5.5.0 ships that code, the issue will remain.

Why it may have gone unnoticed

Nothing in the repository calls this API. hailortcli run2 uses run_async (hailortcli/run2/network_runner.cpp:590), and the Python and GStreamer bindings do not expose it. It is declared in the public header (libhailort/include/hailo/infer_model.hpp:387-388) without a doc comment, and it is not in the HailoRT 5.4.0 User Guide. My guess: on devices that do not use hRPC (e.g. Hailo-8), ConfiguredInferModelImpl is called directly, so line 648 is never reached.

Side notes (not tested)

  • Confidence medium: after one measurement, the device-side implementation calls shutdown() (infer_model.cpp:946). The host client keeps no state about this, so reusing the same ConfiguredInferModel afterwards will probably fail. Is a new configure() expected for each measurement?

  • infer_model.cpp:954 calls expected_fps.value() even when expected_fps holds an error. This is unrelated to the failure above, but that path reads an invalid value.

  • Confidence low: the completion callback counts frames that end with HAILO_STREAM_ABORT (infer_model.cpp:935-941), and total_frames is read after shutdown(). Frames aborted at shutdown may therefore be counted in fps. I have not checked whether abort calls the callback for in-flight frames.

Questions

  1. Is this a known issue? Is a fix planned, for example in 5.5.0?

  2. Is run_async_for_duration() meant to be usable on Hailo-10H through the public API?

  3. Until a fixed release, is there a recommended way to measure on-chip throughput without per-frame host-device transfers? For now I use hailortcli run2, which includes those transfers.