Root cause, the proposed change, and the fault code it reports
Case LCM-866
Container image pulls on prplOS devices fail with “Operation too slow” when the link is silent for ten seconds.
The failure was first seen on hardware in the test lab in April 2026. It has appeared in QEMU CI since September 2026.
The device reports fault 7002 “Request denied” and does not try again. The controller has to start a new pull.
The same code path runs on hardware and in QEMU CI.
Measured in the lab. Twelve seconds of total loss, abort at 10.01 s, both runs.
Low bandwidth. At 1.5 Mbit/s the longest request took 221 s. The guard never fired.
libcurl. No defect found.
The registry. No 429 and no 5xx on any failed request. Every failed request had no HTTP status at all.
Other CI failure classes exist. They are separate topics.
QEMU x86-64 device, faults injected on the WAN link, the real registry. A is today’s code. B is the proposed change.
| condition on the link | A: today (10 s guard, no retry) | B: proposed (30 s guard, 3 retries) |
|---|---|---|
| no fault | pass, longest request 1.3 s | pass, longest request 1.4 s |
| 1.5 Mbit/s for the whole pull | no abort, longest request 221 s | no abort, longest request 216 s |
| total loss for 12 s | abort at 10.01 s, fault 7002, fail | download pauses, completes at 16 s, pass |
| total loss for 40 s | not run, A fails at 10 s | abort at 30 s, one retry, pull completes |
| B with today’s 10 s guard, loss for 12 s | abort at 10 s, one retry after 1 s, pass |
Single runs. The 12 s and 40 s cases ran twice. The 1.5 Mbit/s and 40 s runs still failed the test harness’s own 30-second call deadline. That deadline is a test problem, not a device problem.
USP and TR-069 say the device must not retry a failed state change on its own. Neither text speaks about a second transport request inside one attempt. Both give the device 24 hours to complete the operation. They do not say what the device can do inside those 24 hours.
USP R-SMM.1: “If a DU state change fails, the device MUST NOT attempt to retry the state change on its own initiative”; “MUST complete the requested operation within 24 hours of responding to the InstallDU(), Update() or Uninstall() command”. TR-069 Amendment 6, A.4.1.10: same no-retry wording, DUStateChangeComplete within 24 hours, each operation within one hour.
| client | what it does |
|---|---|
| moby | 5 attempts per layer |
| containerd | 5 attempts per request |
| podman | 3 retries of the pull |
| crane | 3 attempts per request |
| oras | 5 retries per request |
| skopeo | 0 by default, opt-in |
None of them has a byte-rate guard. containerd gives up after 5 minutes without progress.
Decision 1 on the last slide.
| setting | today | proposed |
|---|---|---|
| low-speed limit | 30 B/s, fixed in code | 30 B/s |
| low-speed time | 10 s, fixed in code | 30 s |
| connect timeout | none, libcurl default 300 s | 10 s |
| retries per request | 0 | 3 |
| first wait | 2 s | |
| wait cap | 8 s | |
| wait multiplier | 2 |
Retried, after the wait: operation timed out, could not connect, could not resolve, receive error, send error, partial file, got nothing.
Not retried: HTTP 4xx and 5xx, including 429. TLS errors. Disk errors. Hash and signature errors.
Waits are random, 0 to 2 s, then 0 to 4 s, then 0 to 8 s. 14 s at most per request. After the last attempt the fault code is 7002, as today. The reply and the event do not change. The controller cannot read or set the settings.
Retry-After. A 429 from the registry is a final failure.For InstallDU and Update, TR-181 says: if “the server specified in the URL is not currently reachable or the request times out”, the device SHOULD reject the operation with 7033 “Server Unreachable”. No LCM component emits 7033 today.
TR-181 Issue 2 Amendment 21 (USP), Device.
The controller cannot tell a dead link from a wrong URL from a denied token.
| curl result after the last attempt | today | proposed |
|---|---|---|
| operation timed out, could not connect, could not resolve, got nothing | 7002 | 7033 |
| receive error, send error, partial file | 7002 | 7033 |
| TLS error, HTTP status errors, disk error, hash mismatch | 7002 | 7002 |
A controller that sees 7033 can schedule a later pull. A controller that sees 7002 has no reason to try again. The change is one mapping in the downloader’s error path. The event and the data model do not change.
TR-181 names “the request times out”. Whether a low-speed abort after the last retry counts as that timeout is a reading, not a measurement. Proposed: ask the Broadband Forum to confirm it.
| option | what it does | what it costs | state |
|---|---|---|---|
| Skip the 21 tests that pull an image | Marks them skipped before setup. Re-enable when two full runs are green. | Every assertion in those tests stops running, not only the pull. | prepared, not merged |
| Retry inside the test helpers | The install and update helpers repeat the operation up to 3 times, 5 s apart, on a 7002 with a transport diagnostic. | The tests stop measuring what a device does on a transient error. | prepared, not merged |
| A registry closer to CI | A pull-through cache on the runner, or a registry that prpl operates. | Someone operates it. A cache miss still goes upstream. | 2 person-days for the cache, 4 to 7 for a hosted registry |
Required behavior. Does the device retry a transport error inside one install?
Proposed: yes, bounded as in the change. It merges when the 12-second case passes in CI.
Limits and visibility. Seven startup settings, invisible to the controller.
Proposed: accept now. Propose them to the Broadband Forum as parameters, together with a total deadline.
Fault code. 7033 when the server did not answer, 7002 for everything else.
Proposed: yes, and ask the Broadband Forum to confirm the reading.
Meanwhile. Which fallback keeps the test stack moving?
Proposed: retry inside the test helpers until the change merges.
“No decision yet” is a valid answer for each row.
Lab evidence lives in the campaign repository under reports/T1.3-evidence, not public.