FunASR
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP servin
Star Growth
Overview
FunASR is an industrial speech recognition toolkit supporting offline, streaming, and edge deployment with ASR, VAD, punctuation, and speaker diarization pipelines. It offers OpenAI-compatible API serving and MCP server integration for use with Claude/Cursor, LangChain, Dify, and AutoGen agents. The toolkit includes specialized integrations for voice agents like OpenClaw realtime plugins for self-hosted transcription.
Deep Analysis
Provides both OpenAI-compatible API serving and MCP server specifically for AI agent integration alongside comprehensive speech processing pipelines.
⚡ Capabilities
- • speech recognition (ASR)
- • voice activity detection (VAD)
- • punctuation
- • speaker diarization
- • emotion and audio-event tagging
- • OpenAI-compatible API serving
- • MCP server for Claude/Cursor
🔗 Integrations
✓ Best For
- ✓ adding speech recognition to AI agent workflows
- ✓ building voice-enabled agents with transcription
- ✓ self-hosted speech processing for agent systems
✗ Not Ideal For
- ✗ end-user chat applications
- ✗ general-purpose AI writing tools
- ✗ image generation tasks
⚠ Known Limitations
- ⚠ Multilingual support varies by checkpoint
- ⚠ Timestamp support depends on checkpoint and path
- ⚠ Emotion/event tags do not identify speakers
Alternatives
Works with FunASR
Tools that integrate with FunASR, often used together in the same stack.
Compare FunASR
Maintain FunASR?
Show your live rank in your README, or put FunASR in front of every visitor to AgentoolRank.