Optimize Lexer::readString#1948
Conversation
Lexer::readString previously decoded and validated one UTF-8 character at a time via Lexer::readChar(). Replace this with a bulk scan using strcspn() to copy runs of bytes that need no special handling (i.e. anything other than a quote, backslash, or control character) in a single native call, falling back to the original per-character logic only for quotes, line terminators, control characters, and escape sequences. Care is taken to keep $this->position (a decoded-character count, not a byte count) accurate for multi-byte UTF-8 content. TAB (0x09), a legal, unescaped SourceCharacter per the GraphQL spec, is excluded from the stop-byte set so it is scanned through like any other ordinary character instead of being individually inspected. Visitor::visitInParallel previously called extractVisitFn() to resolve the enter/leave callback for the current visitor and node kind on every single node, for every wrapped visitor - an O(nodes x visitors) number of array lookups, even though a callback's identity for a given (visitor, kind, direction) triple never changes during a traversal. Since NodeKind exposes a small, fixed set of kinds, precompute this mapping once per call (O(visitors x kinds)) instead. Additionally, replace func_get_args() plus argument unpacking in the enter/leave dispatcher closures with the five explicit, typed parameters they are always called with (matching Visitor::visit()'s single call site), avoiding func_get_args()'s per-call overhead. Add a regression test for the TAB-handling edge case described above.
There was a problem hiding this comment.
Pull request overview
This PR introduces targeted performance optimizations in two hot-path components of the library: lexing string literals and running multiple validation visitors in a single traversal.
Changes:
- Optimize
Lexer::readString()by bulk-scanning “ordinary” bytes viastrcspn()and only falling back to per-character handling for escape/control/terminator cases (explicitly allowing unescaped TAB per spec). - Optimize
Visitor::visitInParallel()by precomputing per-(visitor, kind, direction) callbacks once per traversal and avoidingfunc_get_args()/argument unpacking in the dispatcher. - Add a regression test ensuring strings containing an unescaped TAB are lexed correctly.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
| tests/Language/LexerTest.php | Adds regression coverage for unescaped TAB in string literals. |
| src/Language/Visitor.php | Reduces per-node overhead in parallel visitor dispatch by caching callbacks and avoiding func_get_args(). |
| src/Language/Lexer.php | Speeds up string literal lexing by scanning runs of non-special bytes in bulk while preserving escape/control handling. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
The enter/leave callback caches were previously prefilled only for NodeKind constants, silently skipping generic and kind-specific visitors for Nodes with custom (non-NodeKind) kinds. Switch to lazy, per-direction caches that resolve and store a callback (or `false` as a "no callback" sentinel) on first use, keyed by any kind string. Also fixes a pre-existing undefined-array-key warning in visit() when traversing a Node with a custom kind. Add regression tests covering generic and kind-specific visitors for a custom Node kind.
spawnia
left a comment
There was a problem hiding this comment.
Are those two independent changes? Please split if so.
- Use a HEREDOC for the tab-character test source, deriving the expected value from it via substr() so input and expectation can't drift apart. - Type-hint $path/$ancestors as array in visitInParallel()'s enter/leave closures, and document $key/$parent via @phpstan-param since PHP 7.4 doesn't support union types. - Add @var docblocks describing the shape of the enter/leave callback caches. - Replace the lazily-built $stopBytes runtime computation in Lexer::readString() with a precomputed STRING_STOP_BYTES constant.
The visit() undefined-array-key warning fix for custom Node kinds, along with its regression tests, has been split out into a separate PR (webonyx#1952) since it is unrelated to the Lexer/Visitor optimization work here.
|
Thanks for the review! I've pushed two updates:
This PR is now scoped solely to the |
|
I would like the Lexer and Visitor optimizations as separate PRs. |
The Visitor::visitInParallel() enter/leave callback caching optimization has been split out into a separate PR (webonyx#1953), so it can be reviewed independently of the Lexer::readString optimization here.
|
The |
| // null means EOF, will delegate to general handling of unterminated strings | ||
| case null: | ||
| break; |
There was a problem hiding this comment.
Let's add a failing test case first to make the fix durable.
Summary
This PR contains two related performance optimizations for
Lexer::readString()andVisitor::visitInParallel():Lexer::readString(): bulk-scan runs of "boring" bytes withstrcspn()instead of decoding/validating one UTF-8 character at a time viareadChar(). Falls back to the existing per-character logic only for quotes, line terminators, control characters, and escape sequences. TAB (0x09) is excluded from the stop-byte set, since it's a legal, unescapedSourceCharacterthat can simply be scanned through like any other ordinary character.Visitor::visitInParallel():(visitor, kind, direction)once per call, instead of re-resolving it viaextractVisitFn()on every single node for every wrapped visitor.func_get_args()+ argument unpacking in the enter/leave dispatcher closures with the five explicit, typed parameters they are always called with (verified againstVisitor::visit()'s single call site), avoidingfunc_get_args()'s per-call overhead.Why
Parsing and validating GraphQL documents is on the hot path for every request in a GraphQL server.
readString()andvisitInParallel()(used heavily during validation, where many rules are visited together in one pass) are both called very frequently, so their per-call overhead compounds across large schemas/queries.Testing
composer test(1990 tests, 21728 assertions - no failures; the 1 pre-existing warning is from an unrelated, intentionalExecutorLazySchemaTestwarning-assertion test).composer stan(PHPStan) passes with no errors.composer php-cs-fixerpasses with no remaining issues.composer rector(dry-run) reports no changes needed in any file touched by this PR.testLexesStringsWithUnescapedTabCharacter) covering the TAB case mentioned above.Visitor::visitInParallel()'s traversal order/results are unchanged.Benchmarks
Measured with an isolated benchmark harness (not included in this PR) that repeatedly parses and visits a synthetic ~19KB query containing a mix of plain, escaped, and Unicode string literals, 500 iterations each, with Xdebug disabled (PHP 8.3.6, x86_64). Compared directly against the current
upstream/master(3 runs each, averaged):Parser::parse(Lexer)Visitor::visit(parallel)Real-world gains will vary by query/schema shape and string content; this benchmark specifically isolates the two changed functions from surrounding framework overhead. The
Lexer::readString()improvement scales with string literal length, so queries containing long inline string literals (e.g. large text blocks) should see a proportionally larger benefit than this benchmark's average.Risk assessment
substr()/string-concatenation approach as before, just with fewer, larger operations instead of many single-character ones.