Moondream
Moondream (https://moondream.ai/c/playground and documentation: https://huggingface.co/vikhyatk/moondream2) reads and detects elements of an image (Visual Language Model).
In the “playground” area, four modes are offered:
- Query to question the model,
- Caption for a description of the image in three formats,
- Point to ask the model to point to an element,
- Detect to make the model identify one or more elements.
In English with French accepted in queries. For “Point” and “Detect”, best results with English queries.
Unlimited and open-source.
Images for testing:
- A man wearing face mask cycling by Wat Phra Kaew outside the wall of the Grand Palace, Bangkok.
- Örebro slott May
- Giuseppe Arcimboldo – Summer












(Note: difficulties in locating facial parts)