FunASR
Verified · 24 days agoTranscribe local audio with FunASR and SenseVoice using private, on-device inference.
Generate images (Nano Banana), video (Veo, Gemini Omni), speech and music (Lyria) with Google models
claude mcp add gemini-media-mcp -- docker run -i --rm ghcr.io/mordor-forge/gemini-media-mcp:1.0.0
Gemini Media wraps Google media-generation models for images, video, speech, and music through MCP. It is for creators or product teams who want agents to generate media assets from Google model APIs rather than using a standalone UI. The project is very early by stars, and because it touches paid model APIs and media rights questions, users should verify credentials handling, cost controls, and licensing guidance before relying on it.
Transcribe local audio with FunASR and SenseVoice using private, on-device inference.
Give your coding agent access to your Figma data. Implement designs in any framework in one-shot.
Turn long videos into viral vertical shorts and publish them to TikTok, Instagram and YouTube.
Trim, convert, resize, compress, and remix audio and video.
Screen recording, meeting notes, and voice dictation - all with AI
Let any LLM watch a video locally — and search everything it has ever watched.