Skip to content

Restore Ropener stall handling lost in the valar-core rewire (v2.7.0 regression) - #34

Merged
daniel-frenkel merged 3 commits into
mainfrom
fix/on-stall-homing-regression
Jul 21, 2026
Merged

daniel-frenkel merged 3 commits into
mainfrom
fix/on-stall-homing-regression

Conversation

@daniel-frenkel

Copy link
Copy Markdown
Member

Fixes the homing regression shipped in v2.7.0. That release has been demoted
to pre-release and releases/latest now resolves to v2.6.4, so the OTA
button and the valar-flasher are serving known-good firmware while this lands.

What broke

v2.7.0 rewired the Ropener onto valar-core as a remote package. The entity
list came through clean — nothing dropped, renamed, or retyped — so it shipped.
But the product's on_stall handler had been silently replaced by core's
default, which zeroes the stepper position on every stall.

Two live bugs:

1. The device wedged in HOMING. start_homing sets global_state = 3 and
nothing cleared it. After any home the State sensor read HOMING forever, the
cover stayed CLOSING, every button gesture died (all guard on state == 0)
and the schedule refused to run (guards on state != 3). The 200 ms completion
poll only handles states 1/2, so it could not recover either. Only the manual
"Start-Stop Homing" button got the unit back.

2. Any stall corrupted the position reference. A curtain snagging mid-travel
silently redefined that point as home, poisoning the cover percentage and the
persisted position. v2.6.4 only zeroed during homing.

The fix

!extend appends to core's action list rather than replacing it — verified
empirically. Extending alone would have left core's unconditional zeroing firing
before the restored handler, so homing would have looked fixed while the
position corruption continued silently.

!remove drops core's handler outright, so two ordered entries leave exactly
one on_stall:

stepper:
  - id: !extend driver
    on_stall: !remove          # delete core's unconditional zeroing
  - id: !extend driver
    on_stall:                  # install the Ropener's conditional handler

Kept inside the product layer deliberately: no valar-core change, no
valar-motion retag, no repin, and no behaviour change for the generic board
builds (which would have lost their zeroing if core's default became a no-op).
Refactoring core to an explicit on_stall hook remains optional cleanup.

The resolved handler is byte-identical to v2.6.4, including the non-homing
else branch (stop, clear state, publish IDLE, leave position alone).

Also in this PR

  • Restore the Reversed ↺ Motor Direction option. Core had dropped the
    glyph. The string is what Home Assistant automations select by, and existing
    units hold it as persisted state. options is a list, so !extend would have
    yielded three choices — same remove/re-add pattern.
  • fw_version defaults to dev instead of a hardcoded 2.7.0, so local and
    branch builds stop misreporting. Release builds still get the tag injected via
    -s fw_version.
  • tools/regression-gate/ — an entity diff and a behavioural diff of the
    resolved config (on_stall, on_press, script bodies, on_boot, intervals,
    *_action, lambdas), anchored to stable identities so a reordered package
    merge is not noise.

Why the behavioural gate

The entity gate passed v2.7.0 completely clean. Entity names cannot see a
handler being swapped out. The behavioural gate caught, in order:

  1. the on_stall regression itself,
  2. the Motor Direction glyph change (unrelated, also shipped in v2.7.0),
  3. a missing else branch in the first draft of this very fix.

Verification

VAL3100 VAL3000
esphome config clean clean
Entity gate 0 removed, 7 added (intended core diagnostics) identical
Behaviour gate 0 dropped; only the documented scheduling/on_boot refactors identical
on_stall vs v2.6.4 identical identical
OTA asset names + button URL unchanged unchanged

esphome compile produces a full firmware.factory.bin for VAL3100.

The on_boot 600/400/-100 split is a benign layer split: recompute_distance
is pure math, the position restore keeps its relative order, and tz/sun moves
later by design. The scheduling refactor moved the global_state != 3 homing
guard into schedule_open/schedule_close — preserved, not dropped.

Before this ships as v2.7.1 latest

  • Compile VAL3000 (in progress)
  • Bench test on real hardware — flash a VAL3100, home it on a motor, confirm
    homing completes and state clears; confirm a mid-travel stall does not zero
    position. The gate cannot prove StallGuard behaviour on physical hardware.

Units already on v2.7.0 need v2.7.1 to recover; the interim workaround is the
manual "Start-Stop Homing" button.

🤖 Generated with Claude Code

daniel-frenkel and others added 3 commits July 21, 2026 13:19
…rewire

v2.7.0 rewired onto valar-core as a remote package. The entity list came
through clean, so it shipped -- but the product's on_stall handler had been
silently replaced by core's default, which zeroes the stepper position on
EVERY stall.

Two live bugs resulted:

  * The device wedged in HOMING. start_homing sets global_state=3 and nothing
    cleared it, so after any home the State sensor read HOMING forever, the
    cover stayed CLOSING, every button gesture died (all guard on state==0)
    and the schedule refused to run (guards on state!=3). Only the manual
    Start-Stop Homing button recovered it.

  * Any stall corrupted the position reference. A curtain snagging mid-travel
    silently redefined that point as home, poisoning the cover percentage and
    the persisted position.

!extend APPENDS to core's action list rather than replacing it, so extending
alone would have left core's unconditional zeroing firing alongside the fix.
Remove core's handler first, then install the product's -- verified against
the resolved config: exactly one on_stall survives, byte-identical to v2.6.4
including the non-homing else branch.

Also in this change:

  * Restore the "Reversed ↺" Motor Direction option. Core had dropped the
    glyph; the string is what Home Assistant automations select by, and
    existing units hold it as their persisted state.

  * fw_version defaults to "dev" instead of a hardcoded "2.7.0", so local and
    branch builds stop misreporting themselves. Release builds still get the
    real version injected by the workflow via -s fw_version.

  * Add tools/regression-gate -- an entity diff AND a behavioural diff of the
    resolved config (on_stall, on_press, script bodies, on_boot, intervals,
    *_action, lambdas). The entity gate alone passed v2.7.0 clean; the
    behavioural gate is what catches this class of regression, and it is what
    caught the Motor Direction glyph and a missing else branch in the first
    draft of this very fix.

Verified on both boards: esphome config clean, entity gate shows only the 7
intended valar-core diagnostics, behavioural gate shows only the documented
scheduling/on_boot refactors and nothing dropped. VAL3100 compiles to a full
factory.bin. OTA asset names and the GitHub-OTA button URL are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The cover had no device_class, so Home Assistant fell back to generic up/down
arrow controls. A Ropener draws sideways; "curtain" gets the horizontal
open/close controls, the right icon, and better voice-assistant phrasing.

Note this does NOT change the device's own web UI: web_server v3 hardcodes
the cover glyphs ("up", "stop", "down" in render_cover) and never reads
device_class. Home Assistant only.

Both gates clean -- entity list and all automation bodies identical; the
resolved config differs by exactly this one line. Included in v2.7.1 because
the bench-tested build had it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…lient

The API reboot_timeout defaults to 15 minutes: with no *API* client connected
the device reboots. It counts native-API clients only -- a browser sitting on
the web UI is not one.

The Ropener is sold as working without Home Assistant, so any customer driving
it from the browser had a controller that silently rebooted every 15 minutes,
stopping the curtain mid-travel if it happened to be moving at the time.
Observed on the bench unit: "[E][api:127] No clients; rebooting", uptime
resetting on a 15 minute cycle.

Not a rewire regression -- v2.6.4 resolves the same 15min default. It is
pre-existing in every Ropener release; the bench session just made it visible.

Set at the product layer rather than in valar-core: core declares a bare
`api:`, so the option merges in cleanly with no !remove needed, and this ships
without a valar-motion retag. It should move into valar-core later, since the
whole family is sold as HA-optional.

Trade-off: this also removes the watchdog that recovers a wedged API
connection. Acceptable -- the users it was rebooting are precisely the ones not
using the API -- and safe_mode still covers boot loops.

Both gates clean; the resolved config differs by exactly one line
(reboot_timeout 15min -> 0s).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@daniel-frenkel
daniel-frenkel merged commit 0c97dd7 into main Jul 21, 2026
2 checks passed
@daniel-frenkel
daniel-frenkel deleted the fix/on-stall-homing-regression branch July 21, 2026 21:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant