Turn words into pictures and video, right inside your chat. Gemini AI brings Google's Nano Banana Pro (studio-grade image generation and editing) and Gemini Omni Flash (fast text-to-video) straight...
What this app can do
10 tools registered
Generate Video20 tok
Generate a short video with audio from a text prompt using Google's Gemini Omni Flash model. GOOGLE-CONFIRMED PROMPT STRUCTURE (Veo prompt guide): if the user's request is vague, expand it yourself into a fully-specified prompt before calling this tool, covering these elements where relevant: subject: The object, person, animal, or scenery in the video (e.g. cityscape, puppies).; action: What the subject is doing (e.g. walking, running, turning their head).; style: Creative direction via film-style keywords (sci-fi, film noir, cartoon).; camera_positioning_motion: Camera location/movement (aerial view, dolly shot, worms-eye) -- optional.; composition: How the shot is framed (wide shot, close-up, two-shot) -- optional.; focus_lens: Lens effects (shallow focus, macro lens, wide-angle lens) -- optional.; ambiance: Color and light contribution to mood (blue tones, warm tones, night) -- optional.. Audio cues: Dialogue: use quotes for specific speech, e.g. "'This must be the key,' he murmured." | Sound effects (SFX): describe explicitly, e.g. "tires screeching loudly, engine roaring." | Ambient noise: describe the environment's soundscape, e.g. "a faint, eerie hum resonates in the background.". Extra tips: Use descriptive language: adjectives and adverbs paint a clearer picture. | Enhance facial detail by naming 'portrait' as a focus when a face matters..
Generate Image Nano Banana Pro12 tok
Generate or edit an image with Nano Banana Pro (best quality) (gemini-3-pro-image). Premium: highest quality, 4K resolution, advanced text/brand accuracy, up to 5 reference images with high character fidelity. Use this when quality matters most -- complex scenes, accurate text in the image, brand consistency. It is the most expensive of the image tools, so prefer a cheaper one for drafts and iterations. Pass reference_generation_ids to reuse a character/setting from this user's own past generations. GOOGLE-CONFIRMED PROMPT STRUCTURE (image-generation docs): if the user's request is vague, expand it yourself into a fully-specified prompt before calling this tool -- do not pass the user's raw short phrase through untouched. Generation templates: photorealistic_scene: A photorealistic [type of shot] of a [subject description] in a [setting description]. [Description of the light]. Shot from a [camera angle] with a [lens type].; stylized_illustration_sticker: A [style] of a [subject, with details about accessories or actions] doing [activity]. The design features [visual qualities, e.g. bold outlines, cel-shading] and [color/background preference].; accurate_text_in_image: Create a [image type] for [brand/concept] with the text "[text to render]" in a [font style]. The design should be [style description], with a [color scheme].; product_mockup: A high-resolution, studio-lit product photograph of a [product description] on a [background surface/description]. The lighting is a [lighting setup] to [lighting purpose]. The camera angle is a [angle type] to showcase [specific feature]. Ultra-realistic, with sharp focus on [key detail]. [Aspect ratio].; minimalist_negative_space: A minimalist composition featuring a single [subject] positioned in the [bottom-right/top-left/etc.] of the frame. The background is a vast, empty [color] canvas, creating significant negative space. Soft, subtle lighting. [Aspect ratio].; sequential_art_comic: Make a [N] panel comic in a [style]. Put the character in a [type of scene].. Editing templates (when an image is provided): add_remove_element: Using the provided image of [subject], please [add/remove/modify] [element] to/from the scene. Ensure the change is [description of how the change should integrate].; inpainting_semantic_mask: Using the provided image, change only the [specific element] to [new element/description]. Keep everything else in the image exactly the same, preserving the original style, lighting, and composition.; style_transfer: Transform the provided photograph of [subject] into the artistic style of [artist/art style]. Preserve the original composition but render it with [description of stylistic elements].; combine_multiple_images: Create a new image by combining the elements from the provided images. Take the [element from image 1] and place it with/on the [element from image 2]. The final image should be a [description of the final scene].; high_fidelity_detail_preservation: Using the provided images, place [element from image 2] onto [element from image 1]. Ensure that the features of [element from image 1] remain completely unchanged. The added element should [description of how it should integrate].; bring_sketch_to_life: Turn this rough [medium] sketch of a [subject] into a [style description] photo. Keep the [specific features] from the sketch but add [new details/materials].; character_consistency_360: A studio portrait of [person] against [background], [looking forward/in profile looking right/etc.]. Best practices: Be hyper-specific: describe details ("ornate elven plate armor, etched with silver leaf patterns") instead of vague terms ("fantasy armor"). | Provide context and intent: state the image's purpose ("a logo for a high-end, minimalist skincare brand") rather than a bare request ("create a logo"). | Iterate and refine conversationally: don't expect perfection on the first try -- follow up with small changes ("make the lighting warmer"). | Use step-by-step instructions for complex scenes with many elements ("first the background, then the foreground object, finally the highlight"). | Use semantic negative prompts: describe the intended scene positively ("an empty, deserted street") instead of negating ("no cars"). | Control the camera with photographic/cinematic language: wide-angle shot, macro shot, low-angle perspective..
Generate Image Nano Banana 28 tok
Generate or edit an image with Nano Banana 2 (balanced) (gemini-3.1-flash-image). Versatile generalist: strong speed/cost/quality balance, up to 4K, up to 4 reference images for character consistency. The balanced default for most requests: clearly cheaper than Pro while still strong on quality and reference consistency. Pass reference_generation_ids to reuse a character/setting from this user's own past generations. GOOGLE-CONFIRMED PROMPT STRUCTURE (image-generation docs): if the user's request is vague, expand it yourself into a fully-specified prompt before calling this tool -- do not pass the user's raw short phrase through untouched. Generation templates: photorealistic_scene: A photorealistic [type of shot] of a [subject description] in a [setting description]. [Description of the light]. Shot from a [camera angle] with a [lens type].; stylized_illustration_sticker: A [style] of a [subject, with details about accessories or actions] doing [activity]. The design features [visual qualities, e.g. bold outlines, cel-shading] and [color/background preference].; accurate_text_in_image: Create a [image type] for [brand/concept] with the text "[text to render]" in a [font style]. The design should be [style description], with a [color scheme].; product_mockup: A high-resolution, studio-lit product photograph of a [product description] on a [background surface/description]. The lighting is a [lighting setup] to [lighting purpose]. The camera angle is a [angle type] to showcase [specific feature]. Ultra-realistic, with sharp focus on [key detail]. [Aspect ratio].; minimalist_negative_space: A minimalist composition featuring a single [subject] positioned in the [bottom-right/top-left/etc.] of the frame. The background is a vast, empty [color] canvas, creating significant negative space. Soft, subtle lighting. [Aspect ratio].; sequential_art_comic: Make a [N] panel comic in a [style]. Put the character in a [type of scene].. Editing templates (when an image is provided): add_remove_element: Using the provided image of [subject], please [add/remove/modify] [element] to/from the scene. Ensure the change is [description of how the change should integrate].; inpainting_semantic_mask: Using the provided image, change only the [specific element] to [new element/description]. Keep everything else in the image exactly the same, preserving the original style, lighting, and composition.; style_transfer: Transform the provided photograph of [subject] into the artistic style of [artist/art style]. Preserve the original composition but render it with [description of stylistic elements].; combine_multiple_images: Create a new image by combining the elements from the provided images. Take the [element from image 1] and place it with/on the [element from image 2]. The final image should be a [description of the final scene].; high_fidelity_detail_preservation: Using the provided images, place [element from image 2] onto [element from image 1]. Ensure that the features of [element from image 1] remain completely unchanged. The added element should [description of how it should integrate].; bring_sketch_to_life: Turn this rough [medium] sketch of a [subject] into a [style description] photo. Keep the [specific features] from the sketch but add [new details/materials].; character_consistency_360: A studio portrait of [person] against [background], [looking forward/in profile looking right/etc.]. Best practices: Be hyper-specific: describe details ("ornate elven plate armor, etched with silver leaf patterns") instead of vague terms ("fantasy armor"). | Provide context and intent: state the image's purpose ("a logo for a high-end, minimalist skincare brand") rather than a bare request ("create a logo"). | Iterate and refine conversationally: don't expect perfection on the first try -- follow up with small changes ("make the lighting warmer"). | Use step-by-step instructions for complex scenes with many elements ("first the background, then the foreground object, finally the highlight"). | Use semantic negative prompts: describe the intended scene positively ("an empty, deserted street") instead of negating ("no cars"). | Control the camera with photographic/cinematic language: wide-angle shot, macro shot, low-angle perspective..
Generate Image Nano Banana 2 Lite5 tok
Generate or edit an image with Nano Banana 2 Lite (fastest/cheapest) (gemini-3.1-flash-lite-image). Fastest and cheapest option. Not optimized for multiple reference images or multi-turn sequential editing. The cheapest and fastest option -- use it for quick drafts, bulk generation, or when the user says the result need not be perfect. Not suited to multiple reference images or step-by-step editing. Pass reference_generation_ids to reuse a character/setting from this user's own past generations. GOOGLE-CONFIRMED PROMPT STRUCTURE (image-generation docs): if the user's request is vague, expand it yourself into a fully-specified prompt before calling this tool -- do not pass the user's raw short phrase through untouched. Generation templates: photorealistic_scene: A photorealistic [type of shot] of a [subject description] in a [setting description]. [Description of the light]. Shot from a [camera angle] with a [lens type].; stylized_illustration_sticker: A [style] of a [subject, with details about accessories or actions] doing [activity]. The design features [visual qualities, e.g. bold outlines, cel-shading] and [color/background preference].; accurate_text_in_image: Create a [image type] for [brand/concept] with the text "[text to render]" in a [font style]. The design should be [style description], with a [color scheme].; product_mockup: A high-resolution, studio-lit product photograph of a [product description] on a [background surface/description]. The lighting is a [lighting setup] to [lighting purpose]. The camera angle is a [angle type] to showcase [specific feature]. Ultra-realistic, with sharp focus on [key detail]. [Aspect ratio].; minimalist_negative_space: A minimalist composition featuring a single [subject] positioned in the [bottom-right/top-left/etc.] of the frame. The background is a vast, empty [color] canvas, creating significant negative space. Soft, subtle lighting. [Aspect ratio].; sequential_art_comic: Make a [N] panel comic in a [style]. Put the character in a [type of scene].. Editing templates (when an image is provided): add_remove_element: Using the provided image of [subject], please [add/remove/modify] [element] to/from the scene. Ensure the change is [description of how the change should integrate].; inpainting_semantic_mask: Using the provided image, change only the [specific element] to [new element/description]. Keep everything else in the image exactly the same, preserving the original style, lighting, and composition.; style_transfer: Transform the provided photograph of [subject] into the artistic style of [artist/art style]. Preserve the original composition but render it with [description of stylistic elements].; combine_multiple_images: Create a new image by combining the elements from the provided images. Take the [element from image 1] and place it with/on the [element from image 2]. The final image should be a [description of the final scene].; high_fidelity_detail_preservation: Using the provided images, place [element from image 2] onto [element from image 1]. Ensure that the features of [element from image 1] remain completely unchanged. The added element should [description of how it should integrate].; bring_sketch_to_life: Turn this rough [medium] sketch of a [subject] into a [style description] photo. Keep the [specific features] from the sketch but add [new details/materials].; character_consistency_360: A studio portrait of [person] against [background], [looking forward/in profile looking right/etc.]. Best practices: Be hyper-specific: describe details ("ornate elven plate armor, etched with silver leaf patterns") instead of vague terms ("fantasy armor"). | Provide context and intent: state the image's purpose ("a logo for a high-end, minimalist skincare brand") rather than a bare request ("create a logo"). | Iterate and refine conversationally: don't expect perfection on the first try -- follow up with small changes ("make the lighting warmer"). | Use step-by-step instructions for complex scenes with many elements ("first the background, then the foreground object, finally the highlight"). | Use semantic negative prompts: describe the intended scene positively ("an empty, deserted street") instead of negating ("no cars"). | Control the camera with photographic/cinematic language: wide-angle shot, macro shot, low-angle perspective..
Generate Image Nano Banana Legacy5 tok
Generate or edit an image with Nano Banana (legacy) (gemini-2.5-flash-image). Legacy 1024px model. Google recommends Nano Banana 2 Lite instead for new work; kept for compatibility. Legacy 1024px model, kept for compatibility. Google recommends Nano Banana 2 Lite instead for new work -- prefer that unless the user explicitly asks for this one. Pass reference_generation_ids to reuse a character/setting from this user's own past generations. GOOGLE-CONFIRMED PROMPT STRUCTURE (image-generation docs): if the user's request is vague, expand it yourself into a fully-specified prompt before calling this tool -- do not pass the user's raw short phrase through untouched. Generation templates: photorealistic_scene: A photorealistic [type of shot] of a [subject description] in a [setting description]. [Description of the light]. Shot from a [camera angle] with a [lens type].; stylized_illustration_sticker: A [style] of a [subject, with details about accessories or actions] doing [activity]. The design features [visual qualities, e.g. bold outlines, cel-shading] and [color/background preference].; accurate_text_in_image: Create a [image type] for [brand/concept] with the text "[text to render]" in a [font style]. The design should be [style description], with a [color scheme].; product_mockup: A high-resolution, studio-lit product photograph of a [product description] on a [background surface/description]. The lighting is a [lighting setup] to [lighting purpose]. The camera angle is a [angle type] to showcase [specific feature]. Ultra-realistic, with sharp focus on [key detail]. [Aspect ratio].; minimalist_negative_space: A minimalist composition featuring a single [subject] positioned in the [bottom-right/top-left/etc.] of the frame. The background is a vast, empty [color] canvas, creating significant negative space. Soft, subtle lighting. [Aspect ratio].; sequential_art_comic: Make a [N] panel comic in a [style]. Put the character in a [type of scene].. Editing templates (when an image is provided): add_remove_element: Using the provided image of [subject], please [add/remove/modify] [element] to/from the scene. Ensure the change is [description of how the change should integrate].; inpainting_semantic_mask: Using the provided image, change only the [specific element] to [new element/description]. Keep everything else in the image exactly the same, preserving the original style, lighting, and composition.; style_transfer: Transform the provided photograph of [subject] into the artistic style of [artist/art style]. Preserve the original composition but render it with [description of stylistic elements].; combine_multiple_images: Create a new image by combining the elements from the provided images. Take the [element from image 1] and place it with/on the [element from image 2]. The final image should be a [description of the final scene].; high_fidelity_detail_preservation: Using the provided images, place [element from image 2] onto [element from image 1]. Ensure that the features of [element from image 1] remain completely unchanged. The added element should [description of how it should integrate].; bring_sketch_to_life: Turn this rough [medium] sketch of a [subject] into a [style description] photo. Keep the [specific features] from the sketch but add [new details/materials].; character_consistency_360: A studio portrait of [person] against [background], [looking forward/in profile looking right/etc.]. Best practices: Be hyper-specific: describe details ("ornate elven plate armor, etched with silver leaf patterns") instead of vague terms ("fantasy armor"). | Provide context and intent: state the image's purpose ("a logo for a high-end, minimalist skincare brand") rather than a bare request ("create a logo"). | Iterate and refine conversationally: don't expect perfection on the first try -- follow up with small changes ("make the lighting warmer"). | Use step-by-step instructions for complex scenes with many elements ("first the background, then the foreground object, finally the highlight"). | Use semantic negative prompts: describe the intended scene positively ("an empty, deserted street") instead of negating ("no cars"). | Control the camera with photographic/cinematic language: wide-angle shot, macro shot, low-angle perspective..
Check Gemini Connection1 tok
Check whether a Gemini API key is configured and whether the Gemini API is reachable.
Save Gemini Api KeyFree
Save the user's own Gemini API key so their generations can run. ``gemini_api_key`` is scoped per-user, not per-app, so each user brings their own -- this is what the left panel's inline key field submits to.
Get Prompt Guide1 tok
Fetch Google's official prompting guide for Gemini image/video generation -- the structured templates, editing patterns and best practices transcribed from Google's own developer documentation. ALWAYS call this FIRST whenever the user asks you to WRITE, DRAFT, IMPROVE or REVIEW a generation prompt, even when they do not want anything generated yet: it is the authoritative source, so a prompt written from memory instead is a guess. Read-only and free of side effects -- it never generates an image, so calling it to check the docs is always safe. Use kind='image' for stills (default), kind='video' for Veo, kind='all' for both.
Upload Reference Image1 tok
Store an image the user supplies (photo, screenshot, artwork) so it can be used as a reference for image generation. Returns a generation_id to pass in reference_generation_ids. Use when the user attaches or uploads a picture and wants it used as the basis, style or character reference for a new image.
Diagnose Image Pipeline1 tok
Diagnostic: report why generated images may fail to display. Measures the real runtime — Pillow availability, stored file sizes, download timings and base64 payload sizes for recent generations.