AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference Framework
AllSpark is a multimodal spatio-temporal general intelligence model that integrates ten different modalities into a unified framework. Inspired by the human cognitive system and linguistic philosophy, AllSpark leverages the Language as Reference Framework (LaRF) to balance cohesion and autonomy among diverse modalities. The model demonstrates strong adaptability across various spatio-temporal modalities, including RGB, SAR, multispectral, hyperspectral, graph, trajectory, point cloud, language, code, and table.
We are excited to announce that our paper "AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference Framework" has been accepted by IEEE Transactions on Geoscience and Remote Sensing (TGRS). We thank the reviewers and the community for their valuable feedback and support!
- Multimodal Integration: AllSpark integrates ten different modalities, including one-dimensional (language, code, table), two-dimensional (RGB, SAR, multispectral, hyperspectral, graph, trajectory), and three-dimensional (point cloud) modalities.
- Language as Reference Framework (LaRF): Inspired by human cognition, LaRF uses language as a reference to align and interpret diverse modalities, enabling a unified representation space.
- Few-shot Learning: AllSpark excels in few-shot classification tasks for RGB and point cloud modalities without additional training, surpassing baseline performance by up to 41.82%.
To train and evaluate AllSpark on a specific modality, follow this file: run.sh
To perform few-shot learning on RGB or point cloud modalities, follow this folder: fewshot
AllSpark has been evaluated on the following datasets:
- RGB: NWPU-RESISC45, UC-Merced, WHU-RS19
- Multispectral (MSI): EuroSAT
- Hyperspectral (HSI): Pavia University
- SAR: MSTAR
- Trajectory: ETH-UCY
- Graph: METR-LA
- Point Cloud: ModelNet40, ShapeNet, ScanObjectNN
- Language: IMDB
- Code: CodeSearchNet
- Table: PRSA
Please refer to the respective dataset documentation for more details on how to download and preprocess the data.
AllSpark has demonstrated competitive performance across various modalities. Below are some key results:
- RGB: Top-1 accuracy of 94.85% on NWPU-RESISC45.
- Point Cloud: Top-1 accuracy of 89.21% on ModelNet40.
- Trajectory: ADE of 0.47 on ETH-UCY.
- SAR: Top-1 accuracy of 97.24% on MSTAR.
For more detailed results, please refer to the paper.
If you find AllSpark useful in your research, please consider citing our paper:
@ARTICLE{10830573,
author={Shao, Run and Yang, Cheng and Li, Qiujun and Xu, Linrui and Yang, Xiang and Li, Xian and Li, Mengyao and Zhu, Qing and Zhang, Yongjun and Li, Yansheng and Liu, Yu and Tang, Yong and Liu, Dapeng and Yang, Shizhong and Li, Haifeng},
journal={IEEE Transactions on Geoscience and Remote Sensing},
title={AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference Framework},
year={2025},
volume={63},
number={},
pages={1-20},
keywords={Point cloud compression;Training;Adaptation models;Philosophical considerations;Linguistics;Data models;Spatiotemporal phenomena;Trajectory;Cognitive systems;Synthetic aperture radar;General intelligence model;large language model (LLM);multimodal machine learning;spatiotemporal data},
doi={10.1109/TGRS.2025.3526725}
}For any questions or inquiries, please contact:
- Run Shao: shaorun@csu.edu.cn
