
This ComfyUI workflow leverages Google's Gemini model to demonstrate the capabilities of multimodal AI in generating coherent and contextually relevant text. The workflow utilizes several nodes, including the GeminiNode, which serves as the core processing unit, and the LoadImage node, which allows users to input images that the Gemini model can analyze and interpret. By integrating these components, the workflow showcases how Gemini's advanced reasoning capabilities can be applied to both text and image inputs, providing a seamless experience for generating text based on visual content. Additionally, the PreviewAny and BatchImagesNode nodes facilitate the visualization and management of multiple outputs, making it easier to handle and review generated content.
API
Use this workflow from code
Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.
Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.
// Install (beta)
npm i @comfyorg/sdk
// Run "Google Gemini" (TypeScript)
import { Comfy } from "@comfyorg/sdk";
const client = new Comfy({ apiKey: "comfyui-..." });
// This workflow, exported in API format (see note below)
const wf = await client.workflows.fromFile("api_google_gemini_api.json");
const asset = client.assets.fromFile("input.png");
wf.setInput("2", "image", asset); // LoadImage
const job = await client.run(wf);
await job.getOutputs("20")[0].toFile("output.png"); // SaveTextThe SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).
Comfy Cloud API access requires a plan with an API key.
FAQ













