Skip to product information

n8n Workflow: Convert Image to Text with GROQ LLaVA V1.5

n8n Workflow: Convert Image to Text with GROQ LLaVA V1.5

 (200+Reviews)
Regular price £20.99
Regular price £20.99 Sale price
SAVE Sold out
Instant Digital Download
Unlimited Downloads
Lifetime Access in Your Account
🔥
128+ Sold
Popular with n8n builders
23 people viewing
High interest right now
9 added today
Fast-moving digital product
n8n Workflow: Convert Image to Text with GROQ LLaVA V1.5

n8n Workflow: Convert Image to Text with GROQ LLaVA V1.5

Regular price £20.99
Regular price £20.99 Sale price
SAVE Sold out

Turn any image into text—instantly—with GROQ LLaVA V1.5 inside n8n

This n8n automation workflow lets you send an image in Telegram and receive an AI-generated description (image-to-text) using the GROQ LLaVA V1.5 7B multimodal model—optimized for fast inference and visual understanding.

What this workflow does

  • Telegram Trigger: The workflow listens for incoming messages in Telegram.
  • Image Extraction: When an image is sent, it extracts the image data from the file.
  • Vision Inference via GROQ LLaVA V1.5: The workflow calls the GROQ LLAVA V1.5 7B API to interpret the visual content and generate a description.
  • Result Delivery: The generated text description is returned so users can read what the model “sees.”

Setup (Telegram bot)

  • Open Telegram and search for @BotFather.
  • Start a chat and type /newbot to create a new bot.
  • Follow the prompts to choose a name and receive a unique API token.
  • Save your access token and username, then use the workflow to begin sending images for descriptions.

Use cases

  • Extract useful text-style descriptions from screenshots, diagrams, or UI images.
  • Generate captions or summaries for images shared by team members in Telegram.
  • Quick visual documentation for SaaS ops—send an image, get structured understanding in return.

Technical details

  • Nodes/Tech stack: set, telegram, sticky note, http request, extract from file, telegram trigger.
  • AI model: GROQ LLaVA V1.5 7B multimodal vision model for image understanding.
  • Output: image-to-text descriptions generated from the provided image.
View full details