nielsr HF Staff commited on
Commit
f11f4a4
·
verified ·
1 Parent(s): 6e28ec6

Add metadata, paper/project links and usage instructions

Browse files

Hi, I'm Niels, part of the community science team at Hugging Face. I've updated the dataset card to include:
- Relevant `task_categories` (`image-text-to-text`) and `language` tags.
- Links to the paper, project page, and GitHub repository.
- A sample usage section showing how to download the dataset via the CLI as documented in the official repository.
- The BibTeX citation for the paper.

Files changed (1) hide show
  1. README.md +39 -3
README.md CHANGED
@@ -36,10 +36,46 @@ configs:
36
  data_files:
37
  - split: train
38
  path: data/train-*
 
 
 
 
 
 
 
 
39
  ---
40
 
 
41
 
42
- # Dataset Name
43
 
44
- This dataset is associated with the paper:
45
- [SketchVLM: Vision Language Models Can Annotate Images to Explain Thoughts and Guide Users](https://arxiv.org/abs/2604.22875)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  data_files:
37
  - split: train
38
  path: data/train-*
39
+ task_categories:
40
+ - image-text-to-text
41
+ language:
42
+ - en
43
+ tags:
44
+ - visual-reasoning
45
+ - physics-prediction
46
+ - sketching
47
  ---
48
 
49
+ # SketchVLM: Physics Ball Drop Dataset
50
 
51
+ This dataset is part of the **SketchVLM** project, introduced in the paper: [SketchVLM: Vision language models can annotate images to explain thoughts and guide users](https://arxiv.org/abs/2604.22875).
52
 
53
+ [**Project Page**](https://sketchvlm.github.io/) | [**GitHub**](https://github.com/Brandon-Collins7/sketchvlm) | [**Interactive Demo**](https://sketch-vlm-demo.vercel.app/)
54
+
55
+ ## Description
56
+
57
+ SketchVLM is a training-free, model-agnostic framework that enables vision-language models (VLMs) to produce non-destructive, editable SVG overlays on input images to visually explain their reasoning.
58
+
59
+ The **Physics Ball Drop** dataset (based on PHYRE) is one of the benchmarks developed to evaluate visual reasoning. In this task, models must predict the trajectory of a ball through various obstacles and provide a visual explanation of the path.
60
+
61
+ ## Usage
62
+
63
+ As documented in the official GitHub repository, you can download the images and metadata locally using the Hugging Face CLI:
64
+
65
+ ```bash
66
+ huggingface-cli download loganbolton/sketchvlm-physics-ball-drop --repo-type dataset --local-dir datasets/ball_drop
67
+ ```
68
+
69
+ ## Citation
70
+
71
+ ```bibtex
72
+ @misc{collins2026sketchvlmvisionlanguagemodels,
73
+ title={SketchVLM: Vision language models can annotate images to explain thoughts and guide users},
74
+ author={Brandon Collins and Logan Bolton and Hung Huy Nguyen and Mohammad Reza Taesiri and Trung Bui and Anh Totti Nguyen},
75
+ year={2026},
76
+ eprint={2604.22875},
77
+ archivePrefix={arXiv},
78
+ primaryClass={cs.CV},
79
+ url={https://arxiv.org/abs/2604.22875},
80
+ }
81
+ ```