For those who don’t know the tool: ComfyUI is a node-based interface and inference engine for generative AI. It exposes the processing steps as an editable graph: load a model, prompt, run a sampler and save the result. Depending on the models and nodes involved, workflows can generate images, video, audio or 3D content. Artists can inspect connections, change individual settings and save the graph as a reusable JSON file. ComfyUI can run locally or through Comfy Cloud. Think visual programming with a growing appetite for GPU time. Comfy Agent adds a conversational assistant to help build and edit those workflows.
The AI now has an AI assistant. Apparently, even artificial intelligence would rather delegate the node wiring. Comfy Agent brings “conversation” (e.a., bellowing into a microphone, letting it transcribe, and then copy-pasting the Orders ) into the workflow editor: ask for an image setup, a different processing stage or help with an error, and the assistant works on the graph itself. The artist still gets the wonderfully human job of deciding whether the result is any good.
Sidenote: The temptation to use AI to write about an AI that helps you use AI was strong. Triple inception, with the spinning top replaced by a node that still needs connecting. But CatGPT said this would, as always, be a bad idea.

The spaghetti stays editable
According to the launch announcement, you can reference another workflow, drag in an asset or identify a node while continuing to edit. Suggested tasks include comparing video models and applying an existing product-shot workflow to twenty images. Up to five chats retain separate histories and context. Five conversations don’t guarantee five simultaneous GPU jobs; even the assistants have to respect the queue. And we all know how efficient five simultaneous conversations are. By the way, this is your reminder: Get earplugs for the Christmas dinner with the family!
The getting-started guide separates construction from execution: choose a blank workflow in the Agent panel, request an image graph with an output node, and inspect its connections and required inputs before running it. Missing models or files still need supplying or replacing. Natural language is a nicer front door; it does not conjure the missing checkpoint out of office stationery and chewing gum.
The workflow selector fixes the edit target for a request. Switching tabs during a response does not redirect the agent, and it cannot create or switch tabs itself. Give it a new selected target for the next request. Otherwise, the enthusiastic assistant may be doing exactly what you asked, just not where your eyes have wandered.
Ask before spending, not before editing
The run controls distinguish Ask, which approves each workflow execution, from Auto, which permits runs without individual confirmation. Both allow graph edits. Stop ends the current agent response; an already submitted workflow uses separate run and queue controls. The big friendly stop button is not a universal off switch for everything currently eating credits.
Teach it the house rules
Personal skills are saved instructions with a name and a description of when to use them. They can specify dimensions, backgrounds, output nodes and validation habits, and load by name or a matching request. Changes take effect on the next message. Public skills are also available; a pasted SKILL.md can become a personal skill. Installing a skill in an external coding assistant does not install it here. A house style can now come with operating instructions, which is more than most house styles manage.
Local is still on the guest list
The product roadmap still lists Desktop support, local custom-node creation, model/node downloads, GPU-aware recommendations and LLM selection as forthcoming. Cloud currently uses supported Cloud nodes. Anthropic supplies the assistant models; this says nothing about which image or video model your graph must use.
For the planned Desktop route, the data-flow documentation describes local files and chat history, but messages and relevant workflow information still pass through Comfy’s online service. Local workflow nodes use local hardware; partner nodes call their providers. Cloud stores its conversations, workflows and assets centrally, and the two chat histories do not synchronise. Local execution is therefore not a promise that the assistant works offline.
Comfy MCP already connects an external agent to Cloud or local ComfyUI. That is a separate integration: it gives your chosen client tools to submit workflows and retrieve outputs. Its local availability does not mean the built-in Comfy Agent has shipped on Desktop. Similar names, different doors.
An assistant with an expense account
The Cloud pricing page lists RTX PRO 6000 Blackwell hardware with 96 GB VRAM. Standard and Creator jobs have a 30-minute ceiling; Pro extends it to an hour. Creator and higher tiers permit model/LoRA imports. Cloud GPU charges accrue during execution, while Agent chat has its own token charges against the same credit pool. Time spent explaining the brief can therefore cost money even before the image model starts work.
Open Comfy Cloud and select Ask Comfy Agent. Agent access includes initial free tokens, then requires a positive credit balance; a subscription is not necessarily required for the chat itself. The published Opus 4.8 rates per 1,000 tokens are 1.055 credits input, 5.275 output, 1.319 for a five-minute cache write, 2.110 for a one-hour cache write and 0.106 for a cache hit. Generation charges are separate. Cloud subscriptions are US$20/35/100 monthly for Standard/Creator/Pro; yearly billing is US$192/336/960 respectively. Rates checked on 7 October 2026.
| Product / developer | Comfy Agent / Comfy |
| Release status | Publicly available on Cloud; described as beta |
| Host / deployment | Comfy Cloud now; Desktop integration coming soon |
| LLM | Anthropic; token table lists Opus 4.8 |
| Access / charging | Free initial chat tokens, then usage-based Comfy Credits |
| Cloud plans | Standard US$20, Creator US$35, Pro US$100 per month on monthly billing |