Image-Text-to-Text
Transformers
Safetensors
vlm_with_timeseries
visual-question-answering
time-series
multimodal
qwen3-vl
lora
anomaly-reasoning
arfbench
observability
conversational
Instructions to use Datadog/Toto-1.0-QA-Experimental with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Datadog/Toto-1.0-QA-Experimental with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Datadog/Toto-1.0-QA-Experimental") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Datadog/Toto-1.0-QA-Experimental", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Datadog/Toto-1.0-QA-Experimental with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Datadog/Toto-1.0-QA-Experimental" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Datadog/Toto-1.0-QA-Experimental", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Datadog/Toto-1.0-QA-Experimental
- SGLang
How to use Datadog/Toto-1.0-QA-Experimental with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Datadog/Toto-1.0-QA-Experimental" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Datadog/Toto-1.0-QA-Experimental", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Datadog/Toto-1.0-QA-Experimental" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Datadog/Toto-1.0-QA-Experimental", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Datadog/Toto-1.0-QA-Experimental with Docker Model Runner:
docker model run hf.co/Datadog/Toto-1.0-QA-Experimental
|
Download README.md from Datadog/Toto-1.0-QA-Experimental: direct link, hf CLI and curl.
- Browser
- Download file 6.19 kB
-
https://huggingface.co/Datadog/Toto-1.0-QA-Experimental/resolve/main/README.md
- Command line
-
hf download hf://Datadog/Toto-1.0-QA-Experimental/README.md
-
curl -L -o README.md https://huggingface.co/Datadog/Toto-1.0-QA-Experimental/resolve/main/README.md
6.19 kB
| base_model: | |||
| - Qwen/Qwen3-VL-32B-Instruct | |||
| - Datadog/Toto-Open-Base-1.0 | |||
| datasets: | |||
| - Datadog/ARFBench | |||
| license: apache-2.0 | |||
| metrics: | |||
| - accuracy | |||
| - f1 | |||
| pipeline_tag: image-text-to-text | |||
| library_name: transformers | |||
| tags: | |||
| - visual-question-answering | |||
| - time-series | |||
| - multimodal | |||
| - qwen3-vl | |||
| - lora | |||
| - anomaly-reasoning | |||
| - arfbench | |||
| - observability | |||
| model_id: Toto-1.0-QA-Experimental | |||
| leaderboards: | |||
| - ARFBench | |||
| # Toto-1.0-QA-Experimental | |||
| `Toto-1.0-QA-Experimental` is a hybrid time-series foundation model (TSFM) and vision-language model (VLM) for ARFBench. | |||
| The model was introduced in the paper [ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response](https://arxiv.org/abs/2604.21199). | |||
| **Authors:** Stephan Xie, Ben Cohen, Mononito Goswami, Junhong Shen, Emaad Khwaja, Chenghao Liu, David Asker, Othmane Abou-Amal, Ameet Talwalkar. | |||
| ## Model Description | |||
| The model achieves comparable macro F1 and accuracy to top frontier models on ARFBench by combining: | |||
| - A vision-language backbone (`Qwen/Qwen3-VL-32B-Instruct`) for image-conditioned question answering. | |||
| - Toto time-series representations (`Datadog/Toto-Open-Base-1.0`). | |||
| - Lightweight projection modules that inject time-series signals into VLM inference. | |||
| |%7C%3C!----%3E%3C%2Ftd%3E%3C%2Ftr%3E%3Ctr id="L43"> | |:-:| | ||
| |Overall accuracy and F1 on the ARFBench time series question-answering benchmark, as of paper release. Toto-1.0-QA-Experimental achieves the top accuracy and comparable F1 to top frontier models.| | |||
| |%7C%3C!----%3E%3C%2Ftd%3E%3C%2Ftr%3E%3Ctr id="L47"> | |:-:| | ||
| |Overview of the Toto-1.0-QA-Experimental Architecture.| | |||
| This model repository stores inference artifacts, including merged vision-language model weights, time-series modules, and configuration files. | |||
| --- | |||
| ## Basic Inference Example | |||
| The example below assumes you already have time-series tensors, one or more image paths, and a text question. The required components are available in the [official Github repository](https://github.com/DataDog/arfbench). | |||
| ```python | |||
| import torch | |||
| from transformers import AutoProcessor | |||
| from qwen_vl_utils import process_vision_info | |||
| # From our Github repository (https://github.com/DataDog/arfbench) | |||
| from model.toto_vlm_components import TotoAnomalyQAModel, TimeSeriesData | |||
| repo_id = "Datadog/Toto-1.0-QA-Experimental" | |||
| # Load model + processor from Hub artifact | |||
| model = TotoAnomalyQAModel.from_pretrained( | |||
| repo_id, | |||
| device_map="auto", | |||
| torch_dtype=torch.bfloat16, | |||
| ) | |||
| processor = AutoProcessor.from_pretrained(repo_id) | |||
| model.eval() | |||
| # ----------------------------------------------------------------------------- | |||
| # Example input data (replace with your real tensors and inputs) | |||
| # ----------------------------------------------------------------------------- | |||
| series = ... # torch.Tensor, shape: [n_channels, n_timesteps], float32 | |||
| padding_mask = ... # torch.Tensor, shape: [n_channels, n_timesteps], bool | |||
| id_mask = ... # torch.Tensor, shape: [n_channels, n_timesteps], float/bool | |||
| timestamp_seconds = ... # torch.Tensor, shape: [n_channels, n_timesteps] | |||
| time_interval_seconds = ... # torch.Tensor, shape: [n_channels] | |||
| group_names = ... # list[str], length n_channels | |||
| question = "In the following time-series, does the anomaly in this time-series correlate with the anomaly in the other time-series, if anomalies exist??" | |||
| image_paths = ["./image_1.png", "./image_2.png"] | |||
| ts_data = TimeSeriesData( | |||
| series=series, | |||
| padding_mask=padding_mask, | |||
| id_mask=id_mask, | |||
| timestamp_seconds=timestamp_seconds, | |||
| time_interval_seconds=time_interval_seconds, | |||
| num_groups=series.shape[0], | |||
| query_group="custom-query", | |||
| group_names=group_names, | |||
| ) | |||
| # Build multimodal chat input (images + text) | |||
| messages = [ | |||
| { | |||
| "role": "system", | |||
| "content": "You are an expert observability anomaly analyst.", | |||
| }, | |||
| { | |||
| "role": "user", | |||
| "content": ( | |||
| [{"type": "image", "image": p} for p in image_paths] | |||
| + [{"type": "text", "text": question}] | |||
| ), | |||
| }, | |||
| ] | |||
| text_prompt = processor.apply_chat_template( | |||
| messages, | |||
| tokenize=False, | |||
| add_generation_prompt=True, | |||
| ) | |||
| processed_images, _ = process_vision_info(messages) | |||
| inputs = processor( | |||
| text=[text_prompt], | |||
| images=[processed_images], | |||
| return_tensors="pt", | |||
| padding=True, | |||
| ) | |||
| device = next(model.parameters()).device | |||
| inputs = { | |||
| k: v.to(device) if isinstance(v, torch.Tensor) else v | |||
| for k, v in inputs.items() | |||
| } | |||
| # Generate answer | |||
| with torch.no_grad(): | |||
| output_ids = model.generate( | |||
| input_ids=inputs["input_ids"], | |||
| attention_mask=inputs.get("attention_mask"), | |||
| pixel_values=inputs.get("pixel_values"), | |||
| image_grid_thw=inputs.get("image_grid_thw"), | |||
| ts_data=[ts_data], # batch of 1 | |||
| max_new_tokens=512, | |||
| do_sample=False, | |||
| ) | |||
| prompt_len = inputs["input_ids"].shape[1] | |||
| answer = processor.decode( | |||
| output_ids[0, prompt_len:], | |||
| skip_special_tokens=True, | |||
| ).strip() | |||
| print("Answer:", answer) | |||
| ``` | |||
| --- | |||
| ## Minimum Requirements | |||
| Running Toto-1.0-QA-Experimental typically requires multi-GPU setup (tested on 4x A100 40GB). If memory is limited, reduce `--max-ts-length` and/or use quantization flags. | |||
| --- | |||
| ## Resources | |||
| - **Paper:** [ARFBench on ArXiv](https://arxiv.org/abs/2604.21199) | |||
| - **Code:** [GitHub - DataDog/arfbench](https://github.com/DataDog/arfbench) | |||
| - **Dataset:** [Datadog/ARFBench](https://huggingface.co/datasets/Datadog/ARFBench) | |||
| - **Leaderboard:** [ARFBench Space](https://huggingface.co/spaces/Datadog/ARFBench) | |||
| --- | |||
| ## Citation | |||
| ```bibtex | |||
| @misc{xie2026arfbenchbenchmarkingtimeseries, | |||
| title={ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response}, | |||
| author={Stephan Xie and Ben Cohen and Mononito Goswami and Junhong Shen and Emaad Khwaja and Chenghao Liu and David Asker and Othmane Abou-Amal and Ameet Talwalkar}, | |||
| year={2026}, | |||
| eprint={2604.21199}, | |||
| archivePrefix={arXiv}, | |||
| primaryClass={cs.LG}, | |||
| url={https://arxiv.org/abs/2604.21199}, | |||
| } | |||
| ``` |