prpl LCM working group, 1 October 2026

Container pulls that fail on ten seconds of silence

Root cause, the proposed change, and the fault code it reports

Case LCM-866

The case

Container image pulls on prplOS devices fail with “Operation too slow” when the link is silent for ten seconds.

The failure was first seen on hardware in the test lab in April 2026. It has appeared in QEMU CI since September 2026.

The device reports fault 7002 “Request denied” and does not try again. The controller has to start a new pull.

Decide today
  1. Does the device retry a transport error inside one install, or does the controller?
  2. Who sets the limits, and can the controller see them?
  3. Which fault code does a transport timeout report?
  4. What keeps the test stack moving until the change merges?

One image pull is six or more requests, one after another

InstallDU or UpdateOne call. The reply waits until the container runs.pull imageRuns in one worker thread.manifest checkanswers 401get tokentoken serviceget manifestget configget layer 1 ... layer Lredirect, then object storageimage readycreates the containerreply plus event DUStateChange!5 + L requests on one connection handle.One failed request fails the pull.Every layer takes two hops:registry, then object storage.

The same code path runs on hardware and in QEMU CI.

Part 1

The stall

Today: ten seconds of silence end the pull

02468101214 sget layerlayer download in flightno bytes in either directionfewer than 30 bytes per second for 10 scurl 28 “Operation too slow”fault 7002 “Request denied”DUStateChange! 7002Install: unit Failed.Update: old version stays.bytes arrive,nobody listens

Measured in the lab. Twelve seconds of total loss, abort at 10.01 s, both runs.

Three defects on one code path

02468101214 sget layerlayer download in flightno bytes in either directionfewer than 30 bytes per second for 10 scurl 28 “Operation too slow”fault 7002 “Request denied”DUStateChange! 7002Install: unit Failed.Update: old version stays.bytes arrive,nobody listensBBNo connect timeout. Thelibcurl default is 300 seconds.CCNo retry. Every transporterror becomes 7002.AAFixed in code: 30 bytes per second over10 seconds, on every request. No setting.
Not the cause

Low bandwidth. At 1.5 Mbit/s the longest request took 221 s. The guard never fired.

libcurl. No defect found.

The registry. No 429 and no 5xx on any failed request. Every failed request had no HTTP status at all.

In CI the silence hits the layer downloads

  • 11 jobs stalled between 5 and 15 September 2026, all with the same error text.
  • 12 of the 13 stalled requests were layer downloads. One was the token request.
  • Two layers of the test image account for 8 of the 12.
  • One job recovered. Its firmware already had the retry. 30 s of silence, one retry, image pulled 21 s later.
two layers of the test imagetokenlayer5 Sep10 Sep15 Sep

Other CI failure classes exist. They are separate topics.

A twelve-second silence reproduces it. Low bandwidth does not.

QEMU x86-64 device, faults injected on the WAN link, the real registry. A is today’s code. B is the proposed change.

condition on the linkA: today (10 s guard, no retry)B: proposed (30 s guard, 3 retries)
no fault
pass, longest request 1.3 s
pass, longest request 1.4 s
1.5 Mbit/s for the whole pull
no abort, longest request 221 s221 s
no abort, longest request 216 s216 s
total loss for 12 s
abort at 10.01 s, fault 7002, fail
download pauses, completes at 16 s, pass
total loss for 40 s
not run, A fails at 10 s
abort at 30 s, one retry, pull completes
B with today’s 10 s guard, loss for 12 s
abort at 10 s, one retry after 1 s, pass

Single runs. The 12 s and 40 s cases ran twice. The 1.5 Mbit/s and 40 s runs still failed the test harness’s own 30-second call deadline. That deadline is a test problem, not a device problem.

Bug or feature? Two questions decide it

What the standard says

USP and TR-069 say the device must not retry a failed state change on its own. Neither text speaks about a second transport request inside one attempt. Both give the device 24 hours to complete the operation. They do not say what the device can do inside those 24 hours.

USP R-SMM.1: “If a DU state change fails, the device MUST NOT attempt to retry the state change on its own initiative”; “MUST complete the requested operation within 24 hours of responding to the InstallDU(), Update() or Uninstall() command”. TR-069 Amendment 6, A.4.1.10: same no-retry wording, DUStateChangeComplete within 24 hours, each operation within one hour.

What other container clients do
clientwhat it does
moby5 attempts per layer
containerd5 attempts per request
podman3 retries of the pull
crane3 attempts per request
oras5 retries per request
skopeo0 by default, opt-in

None of them has a byte-rate guard. containerd gives up after 5 minutes without progress.

Decision 1 on the last slide.

Where the retry lives decides what it costs

Today: the controller retriesProposed: the downloader retries10 s7002controller must decide:policy, operator, or nextmaintenance windowInstallDUminutes to hoursInstall: the Failed unitmust be removed first.token, manifest, config,every layer again30 swait 2 sreplyseconds, at most 4 attempts and14 s of waiting per request
  • The whole image downloads again.
  • A Failed unit to clean up.
  • Fault 7002 does not say “temporary”, so the controller cannot tell a dead link from a wrong URL.
  • One request repeats.
  • No state change above the downloader.
  • The failure report is delayed.

The proposed change: seven settings, read at startup

settingtodayproposed
low-speed limit30 B/s, fixed in code30 B/s
low-speed time10 s, fixed in code30 s
connect timeoutnone, libcurl default 300 s10 s
retries per request03
first wait2 s
wait cap8 s
wait multiplier2

Retried, after the wait: operation timed out, could not connect, could not resolve, receive error, send error, partial file, got nothing.

Not retried: HTTP 4xx and 5xx, including 429. TLS errors. Disk errors. Hash and signature errors.

Waits are random, 0 to 2 s, then 0 to 4 s, then 0 to 8 s. 14 s at most per request. After the last attempt the fault code is 7002, as today. The reply and the event do not change. The controller cannot read or set the settings.

What the change does not do

  • No resume. A retried layer downloads from byte zero. Measured: the same 2.75 MB again, in 2.7 s.
  • No total deadline. On a dead network the device reports 7002 after about 134 s at worst (4 attempts of 30 s plus 14 s of waiting) instead of 10 s. If the connection itself fails, after about 54 s. A pull that stalls on every request can run for many minutes and still succeed.
  • No Retry-After. A 429 from the registry is a final failure.
  • No controller-visible parameter.
  • The test harness’s own 30-second call deadline still fails tests whose pull succeeds.
Today, dead network: 10 sProposed, dead network: about 134 sderived from the settings, not measured
Part 2

The fault code

Today every transport error becomes 7002 “Request denied”

operation timed outcould not connectcould not resolvegot nothingreceive errorsend errorpartial fileTLS errorHTTP 401 / 404 / 429 / 5xxdisk errorhash mismatch7002Request deniedController
What TR-181 offers

For InstallDU and Update, TR-181 says: if “the server specified in the URL is not currently reachable or the request times out”, the device SHOULD reject the operation with 7033 “Server Unreachable”. No LCM component emits 7033 today.

TR-181 Issue 2 Amendment 21 (USP), Device.SoftwareModules.InstallDU(), FaultCode description.

The controller cannot tell a dead link from a wrong URL from a denied token.

Proposal: report 7033 when the server did not answer

curl result after the last attempttodayproposed
operation timed out, could not connect, could not resolve, got nothing70027033
receive error, send error, partial file70027033
TLS error, HTTP status errors, disk error, hash mismatch70027002

A controller that sees 7033 can schedule a later pull. A controller that sees 7002 has no reason to try again. The change is one mapping in the downloader’s error path. The event and the data model do not change.

TR-181 names “the request times out”. Whether a low-speed abort after the last retry counts as that timeout is a reading, not a measurement. Proposed: ask the Broadband Forum to confirm it.

Part 3

Meanwhile

Three ways to keep the test stack moving until the change merges

optionwhat it doeswhat it costsstate
Skip the 21 tests that pull an imageMarks them skipped before setup. Re-enable when two full runs are green.Every assertion in those tests stops running, not only the pull.prepared, not merged
Retry inside the test helpersThe install and update helpers repeat the operation up to 3 times, 5 s apart, on a 7002 with a transport diagnostic.The tests stop measuring what a device does on a transient error.prepared, not merged
A registry closer to CIA pull-through cache on the runner, or a registry that prpl operates.Someone operates it. A cache miss still goes upstream.2 person-days for the cache, 4 to 7 for a hosted registry

Decisions requested

  1. Required behavior. Does the device retry a transport error inside one install?

    Proposed: yes, bounded as in the change. It merges when the 12-second case passes in CI.

  2. Limits and visibility. Seven startup settings, invisible to the controller.

    Proposed: accept now. Propose them to the Broadband Forum as parameters, together with a total deadline.

  3. Fault code. 7033 when the server did not answer, 7002 for everything else.

    Proposed: yes, and ask the Broadband Forum to confirm the reading.

  4. Meanwhile. Which fallback keeps the test stack moving?

    Proposed: retry inside the test helpers until the change merges.

“No decision yet” is a valid answer for each row.

Found on the way, separate topics

  • A removed unit can stay behind after uninstall and fails the next install. The candidate cause is a lost wake-up in the downloader.
  • The test harness gives every call 30 seconds. Installs are asynchronous, so tests fail while the device is still pulling.
  • The memory cost of the change is not measured.

Where to check

Lab evidence lives in the campaign repository under reports/T1.3-evidence, not public.