STRIVE: 24-frame evaluation explorer

One portable viewer, 20 selected DynSuperCLEVR examples, and all their raw predictions, reference labels, saved evaluator matches, and 24 RGB/segmentation frames. No GPU or model is needed to view this bundle.

Hosted viewer ยท Public dataset / ZIP download

Download and run on your computer

  1. On the hosted viewer, click Download viewer + examples. Alternatively, open the HF repository's Files tab, select visualizer.zip, and download it. Sign in to your HF account first if this repository is private.
  2. Unzip it. Open a terminal in the extracted dyn24-visualizer folder.
  3. Run python3 serve.py (Windows: py serve.py). Python 3.9+ is sufficient; nothing needs to be installed with pip to run the viewer.
  4. Open http://127.0.0.1:8127/ in your browser. Stop with Ctrl+C.

If that port is busy, run python3 serve.py --port 8130 and open http://127.0.0.1:8130/. Do not open index.html directly as a local file: browsers normally block its JSON requests. The local server only reads files from the extracted viewer directory.

You can also use Open another bundle in an already-running viewer to open an extracted compatible folder. Those files stay in your browser: they are not uploaded to Hugging Face or any other server.

Use the viewer

Occlusion is a display rule, not a metric change

GT visibility comes from source instance visibility and is checked against the GT segmentation pixels. For a visible GT object, the cross uses its existing reference PX position. If that position is null, the source instance's projected 2D position is used, so an occluded GT object can still have a cross. Out-of-frame points are not artificially clamped onto the image border.

When a matched GT object is not visible, its predicted circle is hidden as requested. The table still reports whether the prediction emitted a position, and explicitly says the circle was hidden by GT visibility. Unmatched predictions have no GT occlusion identity and use their own predicted PX/flags. No last-visible-point guessing is used. Birth/death mistakes are exposed in the table, not used to silently truncate trajectories.

Sparse tracks use the existing evaluator's forward/back-fill implementation. Coordinates use its 0โ€“1000 convention, mapped to the image's [0, width-1] and [0, height-1] extent. Actual inference-video frames are decoded; source segmentation alignment is checked against the scene's corresponding RGB PNGs. RGB previews are resized to at most 960 px wide; segmentation uses nearest neighbor resize. No new videos are rendered.

Which examples were selected?

This publication uses mixture checkpoint-414's September 4 PX-only predictions, with the September 9 corrected-GT reevaluation, and pools the actual DynSuperCLEVR validation/test splits (99 + 96 clips). It is not the Dyn-adapted Kubric component split. GT captions come verbatim from the new reference labels: semantic categories are retained (e.g. yellow mountain bicycle), with the agreed seven standalone aliases (airliner, minivan, suv, wagon, truck, dirtbike, scooter). Both object assignments and example scores come from the new saved evaluation; the 20 examples are selected again with the unchanged rule below.

Score = mean of temporal/input-object-set/output-object-set F1 at 0.25, 0.5, 0.75, and 1.0: 12 equally weighted numbers.

  1. Retain parseable predictions with strictly positive score.
  2. Compute the linearly interpolated 25th percentile, at (N-1)*0.25.
  3. Failures are the first 10 ascending scores at or above that value.
  4. Successes are the 10 highest scores. They are relative interaction successes, not necessarily correct object captions or flawless scenes.
  5. Ties use scores rounded to 12 decimals, scene ID, then split. The UI retains the original full-precision score. manifest.json records the cutoff, ranks, selections, input checksums, and helper-code checksums.

This visualization rebuild does not rerun inference/evaluation or modify source artifacts. Original local bundles and ZIP archives are retained unchanged.

Feed in other examples

The reusable builder consumes the same four inputs as STRIVE's generation viewer: inference JSONL, dataset JSON, an existing per-sample evaluation JSON, and a dumped-scene directory containing frames, segmentations, and instances. It supports exactly 24 frames and any model with that schema.

Use an existing STRIVE environment with its repository dependencies. Running the already-exported viewer is dependency-free; building new bundles is not. The builder reuses eval.utils.sentinel_parser.parse_sentinel_scene, eval.utils.metrics_utils.build_dense_object, and the repository's segmentation ID remapping. It also needs the existing NumPy, Pillow, and OpenCV packages. It installs nothing and refuses to overwrite an output directory. To avoid importing unrelated model libraries, it loads the exact densifier and its eight pure helper functions from an explicit AST allowlist in the supplied STRIVE checkout. Their source is not rewritten; its checksum is recorded.

Create my-inputs.json, with paths to your own artifacts:

{
  "model": "My model / checkpoint",
  "cells": [{
    "split": "val",
    "dataset_json": "/path/to/val.json",
    "predictions_jsonl": "/path/to/inference.jsonl",
    "report_json": "/path/to/evaluation.json",
    "scene_root": "/path/to/dumped/scenes",
    "instance_filter_mode": "keep_interacting"
  }]
}

Use the original data export's instance_filter_mode (for example, drop_never_visible or keep_interacting). The builder checks mapped captions and object counts and fails if the source mapping disagrees with the targets. For corrected-caption runs, set top-level caption_manifest to the reevaluation's manifest.json (see the repository's selection_config.json). This verifies the exact dataset/prediction hashes and the source instance hashes, then checks the parsed GT captions against that run's recorded names. It never substitutes legacy shortened asset_id strings into the displayed or raw ground truth.

python build.py --repo-root /path/to/strive --config my-inputs.json \
  --output /path/to/new-viewer \
  --scene-id super_clevr_1084 --scene-id super_clevr_1035
python /path/to/new-viewer/serve.py

Repeat --scene-id to select arbitrary examples (including malformed ones, which are shown without repaired prediction overlays). Omit those flags to apply the 10 lower-quartile / 10 highest-score selection. Add multiple cells to pool splits for a single model. JSONL row alignment uses the saved report's 1-based successful_line_numbers, not a newly filtered row order.

Reproducibility and checks

checksums.sha256 covers all exported files except itself and the ZIP. On Linux/macOS, verify an extracted bundle with sha256sum -c checksums.sha256 (or shasum -a 256 -c checksums.sha256 on macOS).

Run lightweight synthetic checks with:

python3 -m unittest discover -s . -p test_visualizer.py -v

These checks do not launch evaluation, GPUs, Slurm jobs, or downloads. publish.py creates a new static HF Space and, optionally, a dataset containing the complete ZIP, selection manifest, and ZIP checksum. Both are private by default; this publication uses --public with the owner's explicit permission. It uploads only the explicit bundle and refuses to overwrite either repository unless --update-existing is explicitly supplied. Updates check both existing targets and their visibility first, pin writes to their current revisions, and keep previous HF revisions and all local archives. No remote files are deleted; only cases in the new manifest are shown by the viewer. The dataset's raw predictions, reference labels, and frames live inside the ZIP.

Publish with an existing write-enabled credential (never paste its contents into the command or commit it). --token-file does not alter the default HF login:

python publish.py --bundle /path/to/exported-viewer \
  --repo-id YOUR_ACCOUNT/NEW_VIEWER \
  --dataset-repo-id YOUR_ACCOUNT/NEW_VIEWER \
  --token-file /path/to/hf-write-token --public

To refresh the existing public viewer and ZIP, rebuild into a new directory with the updated selection_config.json, then run the same publication command with the existing repository IDs and --update-existing --public. Do not use repackage.py for changed scores, labels, or selections: it preserves the old case payloads. A full build.py export is required in that situation.

After code/doc updates, package the same saved cases into a new directory:

python repackage.py --source /path/to/old-viewer --output /path/to/new-release

This verifies source checksums and preserves the old bundle and ZIP. The saved case files and manifest are reused byte-for-byte; no cases are selected again. On the same filesystem they are hardlinked, so treat saved case assets as immutable. The updated code/docs are independent copies.

This viewer is based on STRIVE's eval/serve_kubric_generation_viewer.py and eval/viewer_assets/app.js (synchronized canvases and object controls), eval/serve_oracle_prefill_audit.py (index-based saved match rows), and surya_runs/build_strive_visuals.py (persisted-artifact scoring and packaging).

HF static Space documentation: https://huggingface.co/docs/hub/spaces-sdks-static