Multi-Domain Synthetic Dataset for Rural Driving
Jongwon Ryu, Jaehoon Go, Trung X. Pham, and Junyeong Kim
MIST is a synthetic rural-driving dataset for studying environmental domain variation. This repository accompanies the paper and provides documentation and dataset-access examples. The image data is hosted on Hugging Face, not in GitHub.
This repository was previously named SMS and is now named MIST to match the published paper and dataset.
Unlike urban-centric driving data, MIST focuses on rural roads and variations in background appearance. Scenes are generated with Slowroads and organized into 32 balanced domain configurations:
| Factor | Values |
|---|---|
| Season | Spring, summer, autumn, winter |
| Time of day | Dawn, daytime, dusk, night |
| Weather | Clear, overcast |
The dataset supports research on multi-domain image-to-image translation,
vision-language analysis, and environmental domain shifts. Domain labels are
provided as text, such as autumn dawn clear weather rural road.
The current Hugging Face release provides image-text pairs in Parquet format:
| Split | Image-text pairs |
|---|---|
| Train | 32,000 |
| Test | 3,200 |
| Total | 35,200 |
| Field | Content |
|---|---|
image |
Rural-driving image, decoded as a PIL image when loaded |
text |
Text description of the environmental domain |
The released files occupy approximately 73 GB compressed. Streaming is recommended for inspecting a few samples without downloading the full release. These counts describe the current public subset, not the full dataset described by the paper.
The dataset card lists segmentation annotations and remaining data as ongoing
release work. The current default configuration contains only image and text;
do not assume that segmentation masks are already included. Check the
dataset card and files
for the latest availability.
Use Python 3.10 or newer:
git clone https://github.com/jongwonryu/MIST.git
cd MIST
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
# Save three streamed test images and their domain descriptions.
python examples/preview_dataset.py --split test --limit 3 --output outputs/previewThe requirements pin the tested Datasets version to avoid a reported partial-Parquet-stream shutdown issue in newer scanner releases.
Or use Hugging Face Datasets directly:
from datasets import load_dataset
dataset = load_dataset(
"jongwonryu/MIST-autonomous-driving-dataset",
split="test",
streaming=True,
)
sample = next(iter(dataset))
print(sample["text"])
print(sample["image"].size)
sample["image"].save("mist_sample.png")Image dimensions should be read from the loaded sample rather than assumed from
the simulator's original rendering settings. Set streaming=False only when
you intend to download and cache the requested split locally.
This GitHub release contains dataset documentation, citation metadata, and a small preview utility. It does not currently contain the simulator modifications, dataset-generation pipeline, or training/evaluation code used in the paper.
@article{ryu2026mist,
title = {{MIST}: Multi-Domain Synthetic Dataset for Rural Driving},
author = {Ryu, Jongwon and Go, Jaehoon and Pham, Trung X. and Kim, Junyeong},
journal = {IEEE Access},
volume = {14},
pages = {132866--132877},
year = {2026},
doi = {10.1109/ACCESS.2026.3725755}
}The Hugging Face dataset card declares Apache-2.0 for the dataset. No separate software license has been specified for this GitHub repository. Refer to the dataset card and applicable simulator terms when reusing those materials.
Jongwon Ryu: fbwhddnjs511@cau.ac.kr.
