A hardware decoder cuts client decode power ~3.7× — and having the silicon isn't enough: stock software must be able to reach it, while having none at all can put the codec out of reach entirely
Pi 400, same board, same 1080p60 file — H.264 hw +0.41 W vs sw +1.50 W playing (3.7×, n=6/3); +0.59 vs +2.72 W saturated (4.6×, n=3). Pi 5 (block dropped): +1.57 W.
SCOPE: Client device layer only (Raspberry Pi 400 / Pi 5, headless pure decode; Google TV as playback context). Network, CDN, display excluded on the Pis.
OWL Finding: A hardware decoder cuts client decode power ~3.7× — and having the silicon isn't enough: stock software must be able to reach it, while having none at all can put the codec out of reach entirely
measured 2026-07-29, refined 2026-09-22
https://wattlab.greeningofstreaming.org/findings/hw-decoder-cuts-client-energy-4x
Greening of Streaming — wattlab.greeningofstreaming.org
Source measurement
Loading measurement decode/dec0de04…
Loading measurement decode/dec0de05…
Loading measurement decode/c22219b7…
Loading measurement decode/37374ca2…
Loading measurement decode/bffab9f5…
Loading measurement decode/596c3ee6…
Caveats
Third leg added 2026-09-22 — the case where the silicon is simply absent. The TV Box W5 (Allwinner H618) has hardware H.264/HEVC/VP9 and NO AV1 block. Its three hardware codecs all present 60.0 fps; AV1 falls back to Just Player's in-app libgav1 and presents 1.7 fps against a 1080p60 source — i.e. the software path does not merely cost more, it cannot do the job at all on this class of SoC (4x Cortex-A53). This is the limit case of the finding's own thesis: "software must be able to reach the silicon" becomes "with no silicon, software may not reach real time". n=3, BBB iso-bitrate 1080p60, batch 3e54b322a9b4.
That same absence produces a measurement trap rather than a measurement: because the W5 is not doing the work, its AV1 ΔW (+0.442 W) comes in 58% BELOW the cheapest codec it can actually play (HEVC +0.966 W). Any hardware-vs-software comparison that includes a codec the device cannot sustain will flatter the software path. Frame rate is the only gate that catches it — see CR-078 and docs/w5_onboarding_2026-09-21.md.
Realtime rows are BBB 1080p60 only; single board pair; one rung.
Ratio reconciled 2026-08-09 (R6): n≥3 interleaved under protocol v3 gives 3.7× realtime / 4.6× saturated, replicating July v2 within noise; a 07-30 single-pair read of ~7× rested on one baseline-suspect hw row (+0.221 W, below the n=6 range 0.33–0.51). The hw arm's own rep spread (CV ~18% of ~0.4 W) is why single-pair ratios ranged 3.6–7×.
Pi rows are headless pure decode (ffmpeg -f null, audio disabled) — no display path. A real player adds display/compositor energy on top.
Both Pis' HEVC hardware blocks exist but are unreachable from stock Bookworm userspace (stateless V4L2; GStreamer <1.24, no V4L2-request ffmpeg/mpv, VLC built without the hw paths) — so 'hw vs sw' could only be measured for H.264, on the Pi 400.
Cross-board Pi 400 vs Pi 5 software comparison is n=1 per board with uncontrolled DRAM/clock differences; the same-board hw-vs-sw pair is the clean single-variable read.
The result, in one sentence
On the same Raspberry Pi 400, decoding the same 1080p60 H.264 file, the hardware decode path (bcm2835-codec v4l2m2m) drew +0.41 W (n=6) where software drew +1.50 W (n=3) while playing at 1× — 3.7× less — and the Raspberry Pi 5, which shipped without that hardware block, pays +1.57 W in software for the same stream.
Why this matters
Encode energy is paid once per title; decode energy is paid per viewer, per hour. OWL's bench has measured the server side since day one — this is the first OWL-grade measurement (per-run idle baseline, ΔW, traffic-light confidence, raw samples persisted) of the client layer that REM's field fleet observes, using the local-mW plug principle LEM documents. A device generation that drops a decoder block, or an OS release that can't reach one, multiplies the per-viewer number by ~4 — invisible in any spec sheet, visible at the wall.
The second half of the headline is the sharper finding: on both Pi generations the HEVC hardware sits idle behind software that cannot drive it. Stock-OS users software-decode everything. Stranded silicon is an energy bug.
How it was measured
decode-bench harness (/srv/data/owl/decode-bench/): clips staged in tmpfs (no network, no SD I/O), ffmpeg -f null pure decode (every frame decoded, none displayed), Tapo P110 local mW API at 1.5 s, per-run baseline, OWL confidence.py per row. Two regimes per cell: realtime (-re, 150 s) and full-speed (-stream_loop -1, 120 s). All cited rows 🟢; key realtime rows replicated n=2–3. Full narrative + conjecture list: docs/pi_decode_energy_2026-07.md.
What this finding does not measure
Playback with a display attached (the mpv/KMS arm is future work) — these are decode-only watts.
Any codec other than H.264 for the hw-vs-sw pair (see caveats), any content other than BBB for realtime rows, any resolution other than 1080p.
Whether GStreamer 1.24 unstrands the HEVC blocks (expected, untested — the measured gap predicts roughly 2 W per stream on these boards).