FunASR
Verified · 21 days agoTranscribe local audio with FunASR and SenseVoice using private, on-device inference.
Image, audio and video: generation, transcription, editing, and platform APIs like YouTube or Spotify. Most generation servers proxy a paid API — check whose key you're burning before wiring one into a loop.
Transcribe local audio with FunASR and SenseVoice using private, on-device inference.
Give your coding agent access to your Figma data. Implement designs in any framework in one-shot.
Turn long videos into viral vertical shorts and publish them to TikTok, Instagram and YouTube.
Trim, convert, resize, compress, and remix audio and video.
Screen recording, meeting notes, and voice dictation - all with AI
Let any LLM watch a video locally — and search everything it has ever watched.
The Figma MCP server brings Figma design context directly into your AI workflow.
Natural voice conversations for AI assistants - STT/TTS via MCP
Any file → clean Markdown for AI agents: PDF, Office, EPUB, HTML, images, audio/video. Local MCP.
Extract brand assets (logos, colors, backdrop images, brand name) from any website URL
Edit video in the CartCut desktop editor: cuts from a transcript, captions, motion, effects.
MCP server + Claude Code plugin for ComfyUI: run workflows, generate images, manage models & VRAM.
The AI-powered toolkit that grows your YouTube channel on autopilot
An MCP server retrieving transcripts of YouTube videos
MCP server for Adobe Photoshop — 127 tools (generative AI + recipes), standalone web UI. Control Ph…
An MCP server that provides image generation and editing capabilities
Watch video and live sessions, keep timestamped evidence, and verify an agent's own work.
Render PyTorch architecture diagrams and animated GIF reveals from trusted model source.
Edit, analyze and convert audio: loudness, spec checks, denoise, EQ, cuts, BPM, key. No ffmpeg.
Public Spotify metadata, lyrics and podcasts for LLM agents. No API key, read-only.
MCP server giving AI agents eyes on OpenImageDebugger buffers in live gdb/lldb sessions
Hand any social video to your AI agent — frames + transcript bundled. MCP server for watch-cli.
Bidirectional Figma MCP — AI draws UI on Figma canvas, reads designs back
Parse PDFs, images, doc, docx, ppt, pptx, xls, xlsx, html into Markdown using MinerU API.
Publish videos to TikTok, Reels, Shorts, X, and Facebook through Taisly.
Generate images, video, and audio with Glif's media-generation agent
Drive Google Flow from an agent: Veo video and Imagen image generation
Automate Google NotebookLM — Q&A with citations, audio, video, content generation
Guardrailed video editing for AI agents: FFmpeg, captions, effects, Hyperframes, and receipts.
AI image generation and editing with prompt optimization and quality presets
MCP Server for Video Jungle - Analyze, Search, Generate, and Edit Videos
MCP server to convert Figma designs to Flowbite UI components in Tailwind CSS
MCP server for Draw Things - local AI image generation on Mac
MCP server for OpenAI Images/Videos and Google GenAI (Veo) media generation.
Access Apple Voice Memos on macOS. List, get audio, extract and generate transcripts.
Visual Intelligence Command Center: A Local Computer Vision Engine for Photo Libraries
MCP server for AI-powered image recognition and description using OpenAI vision models.
MCP server for AI-powered image recognition and description using OpenAI vision models.
Local audio transcription using whisper.cpp. Transcribe with OpenAI Whisper models.
FFmpeg video/audio tools: cut, convert, concat, remove silence, and raw commands.
Submit podcast orders, check status, and manage webhooks via Barevalue editing API.
Turn YouTube videos into short-form clips from any AI assistant
Add text or logo watermarks to images via the Markly.cloud API. Batch supported.
Multi-provider media generation — images, video, audio, and transcription via a unified interface
Upload, transform, and deliver images on a global CDN via AI agents.
Sync Lightroom, Figma, Dropbox & Canva assets to WordPress and Shopify via natural language.
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
Brand rules for AI outputs. Validate and rewrite text to match your brand voice.
Analyze videos: extract frames, transcribe audio, generate storyboard breakdowns.
Sync Lightroom, Figma, Dropbox & Canva assets to WordPress and Shopify via natural language.