NEWS

mocomoco Adds Speaker Diarization to its AI Voice Recognition 'mocoVoice API'!

mocomoco Inc. has added a speaker diarization feature to its high-performance voice recognition AI, the "mocoVoice API".

With this new feature, audio data from multi-person conversations and meetings can be separated and transcribed by individual speaker.

<Features of the New Function>

High-Performance Speaker Diarization
Even in transcriptions involving multiple people, you can clearly identify "who" said "what."

High-Speed Speaker Diarization
Despite the addition of the speaker diarization feature, transcription speed remains unchanged, capable of transcribing 1 hour of audio in just 3 minutes at its fastest.

Multilingual Support

Speaker diarization can be performed with high accuracy even in conversations where Japanese and English are mixed.

<Example Use Cases>

  • Creating minutes for group discussions

  • Recording meetings with business partners involving multiple companies

  • Transcribing events featuring multiple speakers

<About the mocoVoice API>

The mocoVoice API is based on OpenAI Whisper, boasting the highest performance in the voice recognition industry, combined with mocomoco's proprietary dictionary algorithm and acceleration technology. It features the following:

Overwhelming Processing Speed

It can transcribe 1 hour of audio in as little as 3 minutes. Quick transcription is possible even for lengthy meetings and lectures.

Proprietary Dictionary Function

With a dictionary function that requires no phonetic readings, technical terms and proper nouns are recognized accurately. Dictionary registration is available in both Japanese and English.

High-Quality Proofreading by ChatGPT
Recognized text is automatically proofread and formatted into grammatically accurate, easy-to-read sentences. Proofreading is performed according to the linguistic characteristics of both Japanese and English.

Multimedia Support 
In addition to audio files, voice extraction and recognition from video files are also supported.

Code-Switching Support

Even in conversations mixing Japanese and English, language switches are accurately detected and properly transcribed.

<Pricing Plans>

The speaker diarization feature is included in all plans at no additional cost. For more details on the mocoVoice API pricing, please see here: https://docs.mocomoco.ai/guides/pricing

<Development Background>

In meetings and dialogues involving multiple people, accurately grasping "who spoke" was a significant challenge for information sharing and efficient minute-taking. In conventional transcriptions, the inability to identify speakers often increased workload and compromised communication accuracy. To solve these issues, mocomoco Inc. developed the speaker diarization feature for the "mocoVoice API," which can separate speakers quickly and accurately.

<Future Outlook>

We plan to provide a mocoVoice demo page where you can test this feature.

mocomoco will continue to make improvements so that mocoVoice can be used in situations aligned with real-world experiences.

<How to Apply for Service Use>

To start using the mocoVoice API, please apply via the API Usage Application Form below. You can try the new feature immediately after creating an account.

Related Pages

Contact