This repository is released for TMM (IEEE Transactions on Multimedia) review.
Weight is not released at this moment. The trained checkpoints (
.pth.tar) are intentionally excluded. For speed benchmarking, use random weights (--random_weights 1); timing does not depend on weight values.
Supports the paper claim:
KD-NVC-T (tiny) decodes 1080p P-frames at ≈ 14.1 ms/frame ≈ 70.8 FPS on an NVIDIA RTX 5060.
This is the CUDA inference version: it relies on two compiled extensions
(MLCodec_extensions_cpp for rANS entropy coding and inference_extensions_cuda
for the custom CUDA kernels). Build instructions are in §3.
Included:
src/— model architectures and the source of both extensions (.py,.cu,.cpp,.h), which are required to reproduce the speed.logs/round1_cuda.log— the authoritative CUDA-path evidence (70.8 FPS).logs/round2_torch.log— PyTorch-fallback reference for consistency.test_video.py— inference entry point with--random_weightsfor speed-only runs.build_extensions.sh— one-shot build of both extensions.docs/EVIDENCE.md— environment and measurement methodology.
Not included:
- Trained weights (
checkpoint/*.pth.tar) — weight is not released at this moment. - Prebuilt binaries (
*.so) — build them from source, see §3. - Test video data (
dataset/) — bring your own YUV sequence, see §4.
Encoding/decoding time is determined by the model architecture + CUDA kernels,
not by the numerical values of the weights. --random_weights 1 keeps the randomly
initialized model (it never loads a checkpoint) and runs the full forward pass, so
the measured latency matches the trained model; only PSNR / bitrate are meaningless.
The pipeline depends on two compiled extensions:
| # | Extension | Purpose | Backend |
|---|---|---|---|
| 1 | MLCodec_extensions_cpp |
rANS arithmetic entropy coding | C++ (pybind11) |
| 2 | inference_extensions_cuda |
custom CUDA quantization/conv kernels | CUDA |
- CUDA Toolkit 13.0 (
nvcc, sm_120 for RTX 5060 Blackwell) - Python 3.12 + PyTorch 2.9.1 built against CUDA 13.0
- Build deps:
pip install pybind11 ninja
CUDA_HOME=/usr/local/cuda TORCH_CUDA_ARCH_LIST=12.0 bash build_extensions.shbuild_extensions.sh builds both extensions in place and copies the resulting
.so files into the active interpreter's site-packages.
Project 1 — MLCodec_extensions_cpp (rANS entropy coder, CPU):
cd src/cpp
python setup.py build_ext --inplace
python -c 'import site, glob, shutil, os; \
[shutil.copy(f, site.getsitepackages()[0]) for f in glob.glob("MLCodec_extensions_cpp*.so")]'
cd ../..Project 2 — inference_extensions_cuda (custom CUDA kernels):
cd src/layers/extensions/inference
CUDA_HOME=/usr/local/cuda TORCH_CUDA_ARCH_LIST=12.0 \
python setup.py build_ext --inplace
python -c 'import site, glob, shutil, os; \
[shutil.copy(f, site.getsitepackages()[0]) for f in glob.glob("inference_extensions_cuda*.so")]'
cd ../../../..python - <<'PY'
import MLCodec_extensions_cpp
import inference_extensions_cuda
from src.layers.cuda_inference import CUSTOMIZED_CUDA_INFERENCE
print("extensions: OK")
print("CUSTOMIZED_CUDA_INFERENCE =", CUSTOMIZED_CUDA_INFERENCE) # must be True
PYIf CUSTOMIZED_CUDA_INFERENCE is False, the code silently falls back to pure
PyTorch (correct but slower); rebuild the CUDA extension for your toolchain.
mkdir -p dataset
# Place any 1080p yuv420 8-bit sequence (>= 20 frames) at:
# dataset/Beauty_1920x1080_120fps_420_8bit_YUV.yuv
# or edit dataset_config_single1080p_yuv420.json accordingly.bash test.shEquivalent command:
python test_video.py \
--model_type tiny \
--test_config dataset_config_single1080p_yuv420.json \
--output_path output_tiny_cuda_1080p_qp32.json \
--qp 32 --force_frame_num 20 --random_weights 1 \
--cuda 1 --write_stream 1 --verbose 2 --stream_path out_bin_cudaExpected: dec frame 2..19 each ≈ 13.7–14.7 ms; analyze_log.py reports
steady-state P-frame decoding ≈ 70.8 FPS.
logs/round1_cuda.log reports Dec avg : 17.1 ms/f (58.6 FPS), which is the
average over all 20 frames including the I-frame and the first warm-up P-frame.
The paper metric excludes warm-up and counts only steady-state P-frames:
| Metric | Frames | Avg latency | FPS |
|---|---|---|---|
| Raw log average | all 20 | 17.1 ms/f | 58.6 |
| P-frames (excl. I-frame) | frame 1–19 | ~14.3 ms/f | ~70.0 |
| Steady-state P-frames (paper) | frame 2–19 | ≈ 14.1 ms/f | ≈ 70.8 |
The I-frame decode (70.07 ms) and the first P-frame (16.96 ms, warm-up / adaptive-I preparation) are excluded. Reproduce the table with:
python analyze_log.py logs/round1_cuda.logSee docs/EVIDENCE.md for the full environment and methodology.
.
├── src/
│ ├── models/ # image / video_model / tiny / small / standard
│ ├── layers/ # cuda_inference, layers, CUDA ext source (.cu/.cpp)
│ ├── cpp/ # rANS entropy coder source (pybind11)
│ └── utils/ # YUV I/O, PSNR, bitstream helpers
├── logs/
│ ├── round1_cuda.log # CUDA path (authoritative)
│ └── round2_torch.log # PyTorch fallback reference
├── docs/EVIDENCE.md
├── test_video.py # --random_weights for speed-only runs
├── compare_consistency.py # CUDA vs PyTorch consistency check (optional)
├── analyze_log.py # steady-state P-frame FPS from a log
├── test.sh # one-shot benchmark
├── build_extensions.sh # build both extensions
├── dataset_config_single1080p_yuv420.json
└── requirements.txt
conda create -n kdnvc python=3.12 -y
conda activate kdnvc
pip install torch==2.9.1 --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt # numpy scipy pillow
pip install pybind11 ninja # build-time onlyAdapted from Microsoft DCVC (MIT). Released under the MIT License, see LICENSE.