Skip to content

refactor: measure splitTime in doSplit and shuffleWriteTime at operator level#783

Open
HaoChen-ch wants to merge 2 commits into
bytedance:mainfrom
HaoChen-ch:shuffle_write_metrics
Open

refactor: measure splitTime in doSplit and shuffleWriteTime at operator level#783
HaoChen-ch wants to merge 2 commits into
bytedance:mainfrom
HaoChen-ch:shuffle_write_metrics

Conversation

@HaoChen-ch

@HaoChen-ch HaoChen-ch commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Type of Change

  • 🐛 Bug fix (non-breaking change which fixes an issue)
  • ✨ New feature (non-breaking change which adds functionality)
  • 🚀 Performance improvement (optimization)
  • ⚠️ Breaking change (fix or feature that would cause existing functionality to change)
  • 🔨 Refactoring (no logic changes)
  • 🔧 Build/CI or Infrastructure changes
  • 📝 Documentation only

Description

  1. Measure splitTime directly. Replace the subtraction with a dedicated NanosecondTimer splitTimer(&splitTime_) scoped to the real split work inside doSplit() (buildPartition2Row + preAlloc + splitRowVector). finalizeMetrics() now simply assigns metrics_.splitTime= splitTime_.

  2. Measure shuffleWriteTime at the operator level. Move it into SparkShuffleWriter (ShuffleWriterNode), wrapping addInput (split), reclaim (spill) and noMoreInput→stop. It now reflects the full wall time of the shuffle-write operator (init + addInput + reclaim + stop) instead of just split + stop.

  3. Narrow the flattenTime scope. Tighten the NanosecondTimer(&flattenTime_) blocks so they cover only the flatten calls (ensureFlatten / ensurePartialFlatten / ensureVectorLoaded), excluding the subsequent split/evict logic.

  4. Remove the now-unused totalSplitTime_ and stopTime_ members and their timers in split() / stop() across all three writers.

Performance Impact

  • No Impact: This change does not affect the critical path (e.g., build system, doc, error handling).

  • Positive Impact: I have run benchmarks.

    Click to view Benchmark Results
    Paste your google-benchmark or TPC-H results here.
    Before: 10.5s
    After:   8.2s  (+20%)
    
  • Negative Impact: Explained below (e.g., trade-off for correctness).

Release Note

Please describe the changes in this PR

Release Note:

Release Note:
- Fixed a crash in `substr` when input is null.
- optimized `group by` performance by 20%.

Checklist (For Author)

  • I have added/updated unit tests (ctest).
  • I have verified the code with local build (Release/Debug).
  • I have run clang-format / linters.
  • (Optional) I have run Sanitizers (ASAN/TSAN) locally for complex C++ changes.
  • No need to test or manual test.

Breaking Changes

  • No

  • Yes (Description: ...)

    Click to view Breaking Changes
    Breaking Changes:
    - Description of the breaking change.
    - Possible solutions or workarounds.
    - Any other relevant information.
    

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant