Repository navigation
Restore Ropener stall handling lost in the valar-core rewire (v2.7.0 regression) - #34
Merged
Merged
Conversation
…rewire
v2.7.0 rewired onto valar-core as a remote package. The entity list came
through clean, so it shipped -- but the product's on_stall handler had been
silently replaced by core's default, which zeroes the stepper position on
EVERY stall.
Two live bugs resulted:
* The device wedged in HOMING. start_homing sets global_state=3 and nothing
cleared it, so after any home the State sensor read HOMING forever, the
cover stayed CLOSING, every button gesture died (all guard on state==0)
and the schedule refused to run (guards on state!=3). Only the manual
Start-Stop Homing button recovered it.
* Any stall corrupted the position reference. A curtain snagging mid-travel
silently redefined that point as home, poisoning the cover percentage and
the persisted position.
!extend APPENDS to core's action list rather than replacing it, so extending
alone would have left core's unconditional zeroing firing alongside the fix.
Remove core's handler first, then install the product's -- verified against
the resolved config: exactly one on_stall survives, byte-identical to v2.6.4
including the non-homing else branch.
Also in this change:
* Restore the "Reversed ↺" Motor Direction option. Core had dropped the
glyph; the string is what Home Assistant automations select by, and
existing units hold it as their persisted state.
* fw_version defaults to "dev" instead of a hardcoded "2.7.0", so local and
branch builds stop misreporting themselves. Release builds still get the
real version injected by the workflow via -s fw_version.
* Add tools/regression-gate -- an entity diff AND a behavioural diff of the
resolved config (on_stall, on_press, script bodies, on_boot, intervals,
*_action, lambdas). The entity gate alone passed v2.7.0 clean; the
behavioural gate is what catches this class of regression, and it is what
caught the Motor Direction glyph and a missing else branch in the first
draft of this very fix.
Verified on both boards: esphome config clean, entity gate shows only the 7
intended valar-core diagnostics, behavioural gate shows only the documented
scheduling/on_boot refactors and nothing dropped. VAL3100 compiles to a full
factory.bin. OTA asset names and the GitHub-OTA button URL are unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The cover had no device_class, so Home Assistant fell back to generic up/down
arrow controls. A Ropener draws sideways; "curtain" gets the horizontal
open/close controls, the right icon, and better voice-assistant phrasing.
Note this does NOT change the device's own web UI: web_server v3 hardcodes
the cover glyphs ("up", "stop", "down" in render_cover) and never reads
device_class. Home Assistant only.
Both gates clean -- entity list and all automation bodies identical; the
resolved config differs by exactly this one line. Included in v2.7.1 because
the bench-tested build had it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…lient The API reboot_timeout defaults to 15 minutes: with no *API* client connected the device reboots. It counts native-API clients only -- a browser sitting on the web UI is not one. The Ropener is sold as working without Home Assistant, so any customer driving it from the browser had a controller that silently rebooted every 15 minutes, stopping the curtain mid-travel if it happened to be moving at the time. Observed on the bench unit: "[E][api:127] No clients; rebooting", uptime resetting on a 15 minute cycle. Not a rewire regression -- v2.6.4 resolves the same 15min default. It is pre-existing in every Ropener release; the bench session just made it visible. Set at the product layer rather than in valar-core: core declares a bare `api:`, so the option merges in cleanly with no !remove needed, and this ships without a valar-motion retag. It should move into valar-core later, since the whole family is sold as HA-optional. Trade-off: this also removes the watchdog that recovers a wedged API connection. Acceptable -- the users it was rebooting are precisely the ones not using the API -- and safe_mode still covers boot loops. Both gates clean; the resolved config differs by exactly one line (reboot_timeout 15min -> 0s). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the homing regression shipped in v2.7.0. That release has been demoted
to pre-release and
releases/latestnow resolves to v2.6.4, so the OTAbutton and the valar-flasher are serving known-good firmware while this lands.
What broke
v2.7.0 rewired the Ropener onto
valar-coreas a remote package. The entitylist came through clean — nothing dropped, renamed, or retyped — so it shipped.
But the product's
on_stallhandler had been silently replaced by core'sdefault, which zeroes the stepper position on every stall.
Two live bugs:
1. The device wedged in HOMING.
start_homingsetsglobal_state = 3andnothing cleared it. After any home the State sensor read
HOMINGforever, thecover stayed
CLOSING, every button gesture died (all guard onstate == 0)and the schedule refused to run (guards on
state != 3). The 200 ms completionpoll only handles states 1/2, so it could not recover either. Only the manual
"Start-Stop Homing" button got the unit back.
2. Any stall corrupted the position reference. A curtain snagging mid-travel
silently redefined that point as home, poisoning the cover percentage and the
persisted position. v2.6.4 only zeroed during homing.
The fix
!extendappends to core's action list rather than replacing it — verifiedempirically. Extending alone would have left core's unconditional zeroing firing
before the restored handler, so homing would have looked fixed while the
position corruption continued silently.
!removedrops core's handler outright, so two ordered entries leave exactlyone
on_stall:Kept inside the product layer deliberately: no
valar-corechange, novalar-motionretag, no repin, and no behaviour change for the generic boardbuilds (which would have lost their zeroing if core's default became a no-op).
Refactoring core to an explicit
on_stallhook remains optional cleanup.The resolved handler is byte-identical to v2.6.4, including the non-homing
elsebranch (stop, clear state, publish IDLE, leave position alone).Also in this PR
Reversed ↺Motor Direction option. Core had dropped theglyph. The string is what Home Assistant automations select by, and existing
units hold it as persisted state.
optionsis a list, so!extendwould haveyielded three choices — same remove/re-add pattern.
fw_versiondefaults todevinstead of a hardcoded2.7.0, so local andbranch builds stop misreporting. Release builds still get the tag injected via
-s fw_version.tools/regression-gate/— an entity diff and a behavioural diff of theresolved config (
on_stall,on_press, script bodies,on_boot, intervals,*_action, lambdas), anchored to stable identities so a reordered packagemerge is not noise.
Why the behavioural gate
The entity gate passed v2.7.0 completely clean. Entity names cannot see a
handler being swapped out. The behavioural gate caught, in order:
on_stallregression itself,elsebranch in the first draft of this very fix.Verification
esphome configon_bootrefactorson_stallvs v2.6.4esphome compileproduces a fullfirmware.factory.binfor VAL3100.The
on_boot600/400/-100 split is a benign layer split:recompute_distanceis pure math, the position restore keeps its relative order, and tz/sun moves
later by design. The scheduling refactor moved the
global_state != 3homingguard into
schedule_open/schedule_close— preserved, not dropped.Before this ships as v2.7.1 latest
homing completes and state clears; confirm a mid-travel stall does not zero
position. The gate cannot prove StallGuard behaviour on physical hardware.
Units already on v2.7.0 need v2.7.1 to recover; the interim workaround is the
manual "Start-Stop Homing" button.
🤖 Generated with Claude Code