You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Turn the trace artifacts a run leaves behind into TraceLens operator, kernel, roofline, and collective reports, and add on-demand torch.profiler capture so PyTorch workloads can produce those traces without editing the model script.
Analysis runs in either of two places. On the host, madengine report tracelens keeps TraceLens' pinned protobuf and xprof out of workload images entirely. In-container, the tracelens tool installs into an isolated virtualenv for the same reason. SLURM and Kubernetes collection now gather torch_profiler_output/ and tracelens_output/ per node.
The scripts/ and *.json ignore rules are anchored to the repo root: unanchored they also matched packaged source, which silently excluded the five runtime scripts these tool definitions depend on.
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Integrates TraceLens-based trace analysis into madengine, providing both host-side reporting (madengine report tracelens / tracelens-compare) and in-container tooling to capture PyTorch Kineto traces on-demand (dynolog) and generate TraceLens reports. This fits into the reporting + tools pipeline by turning collected profiler artifacts into actionable operator/kernel/roofline/collective outputs and ensuring distributed collectors gather the needed directories.
Changes:
Add a stdlib-only TraceLens analyzer script plus a host-side wrapper module and new report CLI commands (tracelens, tracelens-compare).
Add in-container tools (torch_profiler_dynolog, tracelens and mode-specific variants) with supporting pre/post scripts and TraceLens-venv isolation.
Extend SLURM/Kubernetes artifact collection, add examples/docs, and add unit/integration/e2e coverage.
Reviewed changes
Copilot reviewed 24 out of 25 changed files in this pull request and generated 3 comments.
Show a summary per file
File
Description
tests/unit/test_tracelens_report.py
Unit tests for host-side TraceLens wrapper functions and report CLI behavior.
tests/unit/test_tracelens_analyze.py
Unit tests for the standalone analyzer script (discovery, command construction, summaries).
tests/integration/test_tracelens_tools_config.py
Integration tests validating tools.json wiring and tool stacking behavior.
tests/e2e/test_tracelens_workflows.py
E2E coverage for host reporting and (GPU) container tool workflows.
The reason will be displayed to describe this comment to others. Learn more.
Two pre-merge nits found while reviewing this PR as the base of the TraceLens stack (#167 -> #170 -> #172 -> #173). Neither is a design issue, both are quick fixes. See inline comments.
Turn the trace artifacts a run leaves behind into TraceLens operator,
kernel, roofline, and collective reports, and add on-demand torch.profiler
capture so PyTorch workloads can produce those traces without editing the
model script.
Analysis runs in either of two places. On the host, `madengine report
tracelens` keeps TraceLens' pinned protobuf and xprof out of workload
images entirely. In-container, the `tracelens` tool installs into an
isolated virtualenv for the same reason. SLURM and Kubernetes collection
now gather torch_profiler_output/ and tracelens_output/ per node.
The `scripts/` and `*.json` ignore rules are anchored to the repo root:
unanchored they also matched packaged source, which silently excluded the
five runtime scripts these tool definitions depend on.
Co-authored-by: Cursor <cursoragent@cursor.com>
…ut a summary
Previously a nonzero analyzer exit with no summary file looked like an empty run
(0 succeeded, 0 failed) and the CLI exited successfully.
Co-Authored-By: Claude <noreply@anthropic.com>
…t of xtrace
DYNOLOG_DEB_URL and TRACELENS_PIP_SPEC were expanded under set -x before the
tracing-suppressed region began, defeating the intended redaction.
Co-Authored-By: Claude <noreply@anthropic.com>
Stale TraceLens summaries can falsely report rerun success
src/madengine/reporting/tracelens_report.py:94
An existing tracelens_summary.json is accepted regardless of the new analyzer's result. If a rerun crashes before writing a summary, the stale file is loaded and its old counters can make the CLI report success. Remove the prior summary before launching the analyzer (or verify that the file was freshly written).
Terminate analyzer option parsing before forwarding extra arguments
Option-like values in extra_args are appended directly to the analyzer's argparse input, so a value such as --detect_recompute is rejected as an unknown analyzer option before TraceLens runs. Terminate this parser's options before forwarding the extra flags.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Turn the trace artifacts a run leaves behind into TraceLens operator, kernel, roofline, and collective reports, and add on-demand torch.profiler capture so PyTorch workloads can produce those traces without editing the model script.
Analysis runs in either of two places. On the host,
madengine report tracelenskeeps TraceLens' pinned protobuf and xprof out of workload images entirely. In-container, thetracelenstool installs into an isolated virtualenv for the same reason. SLURM and Kubernetes collection now gather torch_profiler_output/ and tracelens_output/ per node.The
scripts/and*.jsonignore rules are anchored to the repo root: unanchored they also matched packaged source, which silently excluded the five runtime scripts these tool definitions depend on.