FunASR
Verified · 20 days agoTranscribe local audio with FunASR and SenseVoice using private, on-device inference.
Text-to-speech, speech-to-text, audio-to-face lipsync, and motion-capture clips for 3D agents.
claude mcp add audio-mcp -- npx -y @three-ws/[email protected]
An audio workflow server covering text-to-speech, speech-to-text, lipsync, and motion-capture clips for 3D or embodied agents. It is for developers building avatar, animation, or voice-driven agent experiences rather than ordinary transcription users. The feature set is ambitious, but current adoption signals are modest, so it reads as a specialized project still proving out reliability and depth.
Transcribe local audio with FunASR and SenseVoice using private, on-device inference.
Give your coding agent access to your Figma data. Implement designs in any framework in one-shot.
Turn long videos into viral vertical shorts and publish them to TikTok, Instagram and YouTube.
Trim, convert, resize, compress, and remix audio and video.
Screen recording, meeting notes, and voice dictation - all with AI
Let any LLM watch a video locally — and search everything it has ever watched.