Commit 3ca0534
feat(voice): brevity becomes a rule with a budget, and something that grades it
Brevity was the one headline property of a voice-only interface that the
prompt asserted and nothing measured. It lived as an adjective clause at
block 2 ("one or two sentences"), repeated with no added constraint at
block 4 ("extra short"), competing with 36KB of correctness rules that
mostly push the other way: name the specific action, echo the detail you
heard, account for each outstanding task separately, give the reversal
beat before the result. Length pressure here is structural, and the
prompt never said which side yields.
**A budget, not an adjective.** The new LENGTH block sits after the
hands-free context that motivates it and gives a countable default — one
sentence, under about fifteen words, roughly six seconds — because a
sentence can be forty words and the ear counts time, not punctuation. It
names what to cut, which nothing did before: openers, restating the
request, announcing what it is about to do, unsolicited offers, and a
closing "anything else?" — the call stays open, so it never has to ask.
A worked too-long/right pair carries more than the adjectives did.
**And what it yields to.** LENGTH YIELDS TO ACCURACY, AND TO NOTHING
ELSE. Without that paragraph a tightened brevity rule quietly erodes the
honesty rules the prompt is built around, so the four cases that
genuinely need words are named: a verbatim read-back, two tasks
accounted for separately, presenting a choice's options, why a capture
failed. Blocks 2 and 4 now defer to it instead of stating weaker
versions of it.
**`no-filler` is what makes it hold.** Thirty-eight behaviours in this
prompt hold because a rule grades them; this one did not, and the rubric
had no rule about length or filler at all. The first run proved the
point in the other direction: the rule as first written flagged "I'm
currently checking your unread emails and Slack messages, and after that
I'll book your table" — the exact line the LENGTH block exists to
protect, and a line anyone would be happy to hear on the glasses. The
rule keyed on repeating the request when what matters is whether the
words are an ANSWER. It now says naming a task inside an answer about
that task is content, and carries an operational test for the judge:
flag a line only if words could be deleted with nothing the user asked
for lost. Both rows pass.
Transcripts: a plain calendar check, where nothing needs elaborating and
filler is all there is to add; and `no-filler` added to the existing
"one running, one waiting" row, pointing the other way — that reply
legitimately needs two clauses, so it is where a brevity rule would do
its damage.
Evals (gemini-3.1-flash-lite-preview, judge gemini-3.5-flash-lite):
transcript tier 54/56 effect choice, 74/76 judged, 1 ungraded on a 503;
loop tier 6/6 structural, 8/9 judged. Both `no-filler` rows pass. The
remaining flags are the documented lite-tier ones (`queued-not-underway`
twice, `no-fabricated-timing` once). The two attachLatestImage misses
are NOT this change: A/B'd two runs per prompt on those transcripts and
the pre-change prompt fails the same row the same way.
Counts that had drifted with the catalogue: 41 -> 42 blocks, 31 -> 32
rules, 32 -> 33 transcripts in README, DIRECTORY, SAI_GLASSES_APP and
LoopEvalTest's header.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>1 parent 8e110d5 commit 3ca0534
8 files changed
Lines changed: 39 additions & 11 deletions
File tree
- docs
- meta-android-app/app/src
- main/assets
- test
- java/com/meta/wearable/dat/externalsampleapps/cameraaccess/saispike/eval
- resources/eval
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
91 | 91 | | |
92 | 92 | | |
93 | 93 | | |
94 | | - | |
| 94 | + | |
95 | 95 | | |
96 | 96 | | |
97 | 97 | | |
| |||
100 | 100 | | |
101 | 101 | | |
102 | 102 | | |
103 | | - | |
| 103 | + | |
104 | 104 | | |
105 | 105 | | |
106 | 106 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
206 | 206 | | |
207 | 207 | | |
208 | 208 | | |
209 | | - | |
| 209 | + | |
210 | 210 | | |
211 | 211 | | |
212 | 212 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
363 | 363 | | |
364 | 364 | | |
365 | 365 | | |
366 | | - | |
| 366 | + | |
367 | 367 | | |
368 | 368 | | |
369 | 369 | | |
| |||
Large diffs are not rendered by default.
Lines changed: 2 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
24 | 24 | | |
25 | 25 | | |
26 | 26 | | |
27 | | - | |
28 | | - | |
| 27 | + | |
| 28 | + | |
29 | 29 | | |
30 | 30 | | |
31 | 31 | | |
| |||
Lines changed: 1 addition & 1 deletion
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
13 | 13 | | |
14 | 14 | | |
15 | 15 | | |
16 | | - | |
| 16 | + | |
17 | 17 | | |
18 | 18 | | |
19 | 19 | | |
| |||
Lines changed: 23 additions & 1 deletion
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
326 | 326 | | |
327 | 327 | | |
328 | 328 | | |
| 329 | + | |
| 330 | + | |
| 331 | + | |
| 332 | + | |
| 333 | + | |
| 334 | + | |
| 335 | + | |
| 336 | + | |
| 337 | + | |
| 338 | + | |
| 339 | + | |
| 340 | + | |
| 341 | + | |
| 342 | + | |
329 | 343 | | |
330 | 344 | | |
331 | 345 | | |
| |||
643 | 657 | | |
644 | 658 | | |
645 | 659 | | |
| 660 | + | |
| 661 | + | |
| 662 | + | |
| 663 | + | |
646 | 664 | | |
647 | 665 | | |
648 | | - | |
| 666 | + | |
| 667 | + | |
| 668 | + | |
| 669 | + | |
| 670 | + | |
649 | 671 | | |
650 | 672 | | |
651 | 673 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
155 | 155 | | |
156 | 156 | | |
157 | 157 | | |
| 158 | + | |
| 159 | + | |
| 160 | + | |
| 161 | + | |
| 162 | + | |
158 | 163 | | |
159 | 164 | | |
0 commit comments