Skip to content
Ws-SyxPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

KD-NVC-T · 70.8 FPS 1080p Decoding on RTX 5060 (CUDA inference)

This repository is released for TMM (IEEE Transactions on Multimedia) review.

Weight is not released at this moment. The trained checkpoints (.pth.tar) are intentionally excluded. For speed benchmarking, use random weights (--random_weights 1); timing does not depend on weight values.

Supports the paper claim:

KD-NVC-T (tiny) decodes 1080p P-frames at ≈ 14.1 ms/frame ≈ 70.8 FPS on an NVIDIA RTX 5060.

This is the CUDA inference version: it relies on two compiled extensions (MLCodec_extensions_cpp for rANS entropy coding and inference_extensions_cuda for the custom CUDA kernels). Build instructions are in §3.


1. What is / is not released

Included:

  • src/ — model architectures and the source of both extensions (.py, .cu, .cpp, .h), which are required to reproduce the speed.
  • logs/round1_cuda.log — the authoritative CUDA-path evidence (70.8 FPS).
  • logs/round2_torch.log — PyTorch-fallback reference for consistency.
  • test_video.py — inference entry point with --random_weights for speed-only runs.
  • build_extensions.sh — one-shot build of both extensions.
  • docs/EVIDENCE.md — environment and measurement methodology.

Not included:

  • Trained weights (checkpoint/*.pth.tar) — weight is not released at this moment.
  • Prebuilt binaries (*.so) — build them from source, see §3.
  • Test video data (dataset/) — bring your own YUV sequence, see §4.

2. Why random weights give the same speed

Encoding/decoding time is determined by the model architecture + CUDA kernels, not by the numerical values of the weights. --random_weights 1 keeps the randomly initialized model (it never loads a checkpoint) and runs the full forward pass, so the measured latency matches the trained model; only PSNR / bitrate are meaningless.


3. Build the two extensions (required for CUDA inference)

The pipeline depends on two compiled extensions:

# Extension Purpose Backend
1 MLCodec_extensions_cpp rANS arithmetic entropy coding C++ (pybind11)
2 inference_extensions_cuda custom CUDA quantization/conv kernels CUDA

3.1 Prerequisites

  • CUDA Toolkit 13.0 (nvcc, sm_120 for RTX 5060 Blackwell)
  • Python 3.12 + PyTorch 2.9.1 built against CUDA 13.0
  • Build deps: pip install pybind11 ninja

3.2 One-shot build

CUDA_HOME=/usr/local/cuda TORCH_CUDA_ARCH_LIST=12.0 bash build_extensions.sh

build_extensions.sh builds both extensions in place and copies the resulting .so files into the active interpreter's site-packages.

3.3 Manual build (two projects)

Project 1 — MLCodec_extensions_cpp (rANS entropy coder, CPU):

cd src/cpp
python setup.py build_ext --inplace
python -c 'import site, glob, shutil, os; \
  [shutil.copy(f, site.getsitepackages()[0]) for f in glob.glob("MLCodec_extensions_cpp*.so")]'
cd ../..

Project 2 — inference_extensions_cuda (custom CUDA kernels):

cd src/layers/extensions/inference
CUDA_HOME=/usr/local/cuda TORCH_CUDA_ARCH_LIST=12.0 \
  python setup.py build_ext --inplace
python -c 'import site, glob, shutil, os; \
  [shutil.copy(f, site.getsitepackages()[0]) for f in glob.glob("inference_extensions_cuda*.so")]'
cd ../../../..

3.4 Verify CUDA inference is active

python - <<'PY'
import MLCodec_extensions_cpp
import inference_extensions_cuda
from src.layers.cuda_inference import CUSTOMIZED_CUDA_INFERENCE
print("extensions: OK")
print("CUSTOMIZED_CUDA_INFERENCE =", CUSTOMIZED_CUDA_INFERENCE)  # must be True
PY

If CUSTOMIZED_CUDA_INFERENCE is False, the code silently falls back to pure PyTorch (correct but slower); rebuild the CUDA extension for your toolchain.


4. Prepare test data (not included)

mkdir -p dataset
# Place any 1080p yuv420 8-bit sequence (>= 20 frames) at:
#   dataset/Beauty_1920x1080_120fps_420_8bit_YUV.yuv
# or edit dataset_config_single1080p_yuv420.json accordingly.

5. Run the speed benchmark (random weights, CUDA)

bash test.sh

Equivalent command:

python test_video.py \
  --model_type tiny \
  --test_config dataset_config_single1080p_yuv420.json \
  --output_path output_tiny_cuda_1080p_qp32.json \
  --qp 32 --force_frame_num 20 --random_weights 1 \
  --cuda 1 --write_stream 1 --verbose 2 --stream_path out_bin_cuda

Expected: dec frame 2..19 each ≈ 13.7–14.7 ms; analyze_log.py reports steady-state P-frame decoding ≈ 70.8 FPS.


6. How the 70.8 FPS is computed

logs/round1_cuda.log reports Dec avg : 17.1 ms/f (58.6 FPS), which is the average over all 20 frames including the I-frame and the first warm-up P-frame. The paper metric excludes warm-up and counts only steady-state P-frames:

Metric Frames Avg latency FPS
Raw log average all 20 17.1 ms/f 58.6
P-frames (excl. I-frame) frame 1–19 ~14.3 ms/f ~70.0
Steady-state P-frames (paper) frame 2–19 ≈ 14.1 ms/f ≈ 70.8

The I-frame decode (70.07 ms) and the first P-frame (16.96 ms, warm-up / adaptive-I preparation) are excluded. Reproduce the table with:

python analyze_log.py logs/round1_cuda.log

See docs/EVIDENCE.md for the full environment and methodology.


7. Directory layout

.
├── src/
│   ├── models/                    # image / video_model / tiny / small / standard
│   ├── layers/                    # cuda_inference, layers, CUDA ext source (.cu/.cpp)
│   ├── cpp/                       # rANS entropy coder source (pybind11)
│   └── utils/                     # YUV I/O, PSNR, bitstream helpers
├── logs/
│   ├── round1_cuda.log            # CUDA path (authoritative)
│   └── round2_torch.log           # PyTorch fallback reference
├── docs/EVIDENCE.md
├── test_video.py                  # --random_weights for speed-only runs
├── compare_consistency.py         # CUDA vs PyTorch consistency check (optional)
├── analyze_log.py                 # steady-state P-frame FPS from a log
├── test.sh                        # one-shot benchmark
├── build_extensions.sh            # build both extensions
├── dataset_config_single1080p_yuv420.json
└── requirements.txt

8. Requirements

conda create -n kdnvc python=3.12 -y
conda activate kdnvc
pip install torch==2.9.1 --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt      # numpy scipy pillow
pip install pybind11 ninja           # build-time only

9. License

Adapted from Microsoft DCVC (MIT). Released under the MIT License, see LICENSE.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages