NEWS

Can We Prevent 'He Said, She Said' in Medical Settings? How AI Speech Recognition is Transforming the Frontlines of Pediatric Surgery

This interview article was generated by mocoVoice, the speech recognition AI provided by mocomoco Inc. (our company).

We recently interviewed Dr. Isamu Saeki of Hiroshima University Hospital, who is collaborating with us on the development of the mocoVoice Medical Model.

As AI speech recognition technology evolves rapidly, to what extent is it truly "usable" in fields demanding extremely high expertise and accuracy, such as medical settings? Dr. Saeki, a pediatric surgeon at Hiroshima University Hospital, and Mr. Onishi, in charge of R&D for the "mocoVoice Medical Model" speech recognition AI, discussed the joint development of this medical AI model.


Dr. Isamu Saeki

Entered Kyushu University School of Medicine in 1997 and graduated in 2003. After obtaining his PhD in 2011, he served as Deputy Director and Director of Pediatric Surgery at Fukuoka Children's Hospital and Hiroshima City Hospital. Since 2020, he has been a Lecturer/Clinical Associate Professor of Pediatric Surgery at Hiroshima University Hospital. He holds supervisory physician qualifications from the Japan Surgical Society and the Japanese Society of Pediatric Surgeons, and serves on numerous academic society committees. He has received multiple awards in the areas of clinical practice, education, and research, and has published over a hundred papers in specialized journals.

Kazuyo Onishi, Director & CRO

After completing the regular and advanced courses at Maizuru National College of Technology, he obtained a Master of Engineering from NAIST. He is currently engaged in interactive AI research at RIKEN. He drives business growth by developing products from a user perspective based on meticulous research.


1. The Barrier in Medical Settings: AI Turning "Visual Inspection" into "Guidelines"

Onishi: Thank you for joining us today. Dr. Saeki, you've been collaborating with us on the development of the mocoVoice Medical Model. Could you start by introducing yourself and the background behind this initiative?

Saeki: I'm Saeki, a pediatric surgeon at Hiroshima University Hospital. Pediatric surgeons aren't restricted by organ; we handle all kinds of surgeries for children, from newborns to procedures involving the chest, abdomen, and urinary system.

At the same time, I am deeply committed to the education of medical students. In particular, I am conducting research to build educational programs for training "medical interviewing" and "examination" skills, and I needed AI-powered transcription for this research.

Onishi: That led you to adopt mocoVoice. What were the challenges you faced before that?

Saeki: Initially, I tried using commercially available AI transcription models, but the accuracy was extremely low. Recognizing medical terminology was particularly difficult, and I believe the underlying cause is the abundance of homophones unique to the Japanese language.

For example, doctors refer to "looking" during an examination as "shishin" (視診 - visual inspection). However, general AI models would incorrectly transcribe this as "shishin" (指針 - guidelines). This situation occurred with almost all medical terms.


2. The Power of AI Learning: Correction Time Reduced from "5 Hours" to "Under 1 Hour"

Onishi: The medical field lacks many general terms, and I believe our offering of a specialized medical model contributed to the improved accuracy. Did the situation change with the introduction of mocoVoice?

Saeki: Yes. With regular AI, I had to make more than 10 corrections per page. Even when we first asked mocoVoice and started collaborating to improve accuracy, there were still 6 to 7 corrections per page.

However, this is where I truly experienced the power of AI learning. As we repeated the tuning process nearly 20 times based on over 2 hours of training data, the accuracy improved dramatically. Ultimately, it reached a level where there were only one or two corrections needed per page, or sometimes none at all.

These are actual errors in transcription. While regular transcription has noticeable errors (yellow highlighted sections), tuning the mocoVoice Medical Model clearly reduced the number of errors.

Onishi: Until then, were you manually correcting the mistakes yourself?

Saeki: That's right. Since I had already started the research, I had no choice but to do it myself, but it took about 5 hours each time to listen back to a 2-hour recording and make corrections. Honestly, I was just starting to think I couldn't keep doing it. Towards the end of the research, as mocoVoice's accuracy improved, correcting the same 2-hour recording took less than an hour, which dramatically reduced my burden.


3. Barriers Unique to Pediatric Surgery: "Anterior Fontanelle" and "Poop"

Onishi: In this initiative, we proceeded with tuning for vocabulary specific to your specialty, "pediatric surgery." How was this aspect?

Saeki: We struggled in two areas here. The first was with specialized terms truly only used in pediatrics and pediatric surgery. For example, the "anterior fontanelle" (daisenmon) on a baby's head, the "Apgar score" for evaluating newborns, and conversations about diapers. The AI became able to recognize these as well through repeated learning.

Onishi: I also learned about the "anterior fontanelle" for the first time during the tuning process. It was a classic example of a term where the kanji characters are simple, but recognition is difficult without background knowledge.

Saeki: The other area was words that are generally treated as "vulgar." For instance, words like "poop," "pee," and "wiener." Doctors use these entirely without hesitation during examinations (in fact, it would be bad if we did hesitate).

However, recent AIs, like ChatGPT, tend to filter these words out or act as if they don't exist. Regular transcription software would redact them like "●●●●," which made listening back incredibly tedious.

Onishi: Exactly. General AI is designed not to display obscene or dangerous words. However, this is a "Medical Model." Words spoken by a doctor should be accurately transcribed, even if they are generally considered inappropriate. I made sure to tune the system specifically for this medical field need, intentionally allowing it to recognize and display them.


4. Security and Implementation: Why "On-Premises" is Essential

Onishi: As long as AI is used in healthcare, applications like electronic medical records (EMR) come to mind, but privacy and data security are simultaneously top priorities.

Saeki: That's correct. From a doctor's perspective, my true feeling is that I want the daily burden of typing into EMRs reduced.

Actually, while medical students today are accustomed to typing on smartphones, a surprising number of them cannot touch-type on a computer. Speech recognition that allows for smooth EMR entry, regardless of such skills, is extremely important.

Onishi: It's difficult to upload EMR data to the cloud, isn't it?

Saeki: Yes. In larger hospitals, the network is an "intranet" confined within the hospital. Therefore, an "on-premises" model, where processing is completed on devices within the hospital (on-device) rather than sending data to the internet (cloud), is required.

Onishi: Our company is also considering offering an on-premises solution, but there's a technical trade-off here. The larger and more highly performant the AI model, the more powerful (and expensive) the machine power needed to run it. But if you make the model lighter and smaller, the accuracy drops. With mocoVoice, we are working on the dual goal of "high accuracy, yet running a lightweight and fast model," and we pride ourselves on this being a major strength for on-premises provision.

Actual on-premises provision of mocoVoice (Concept Image)


5. Three Scenarios Medical AI Faces and the Barrier of "Informed Consent"

Saeki: I believe there are broadly three situations where transcription is needed in medical settings.

  1. Conferences: Minutes of meetings where doctors and nurses discuss patient treatment policies.
  2. General Consultations: Assisting with summarizing 1-on-1 conversations between patients and doctors to input into EMRs. Particularly in pediatrics, the basic dynamic is "1-on-2," examining a thrashing child while talking to the parent, so EMR input assistance is a lifesaver.
  3. Informed Consent (IC): The process of explaining treatments in detail to patients and obtaining their consent.

Saeki: This third one is the most important and difficult. In IC, we absolutely cannot have "he said, she said" problems arising later.

Previously, there was a case regarding an endoscopic procedure where a dispute arose over whether it was explained as being "the same as a stomach camera." Here, the AI must not make overly interpreted translations.

Onishi: That's a very interesting point. Even with our AI, because it is a medical model, there are still instances where a patient says they "threw up" (haita), and the AI converts it into the medical term "vomiting" (outo). In a general summary, the meaning comes across, but from an IC perspective, the fact that they specifically said "threw up" absolutely must not be rewritten as "vomiting." This is a crucial challenge for medical AI.

Saeki: Exactly. There were many times when AI development companies brought in what they thought were amazing features saying, "This should be fine," but the honest opinion of doctors was, "Considering IC, we can't use an AI that injects its own interpretations." I feel it is extremely important to have companies like mocoVoice that build products in step with us, based on deep needs in the field (in-depth interviews).


6. In 10 Years, AI Voice Input Will Be "The Norm"

Onishi: Finally, could you share your vision on how the medical model will impact the healthcare industry going forward?

Saeki: Over the past six months, I have strongly felt the benefits of AI. With just about 50 hours of training data, the accuracy improved this much. If this scales to big data, accuracy will increase even further, and I believe we will see a future where EMRs evolve rapidly through a synergistic effect.

When I became a doctor over 20 years ago, we used paper charts. From there, electronic medical records were introduced, and now there are hardly any hospitals without them. This transformation happened in just 10 or 20 years. The times are accelerating, so I imagine that within the next 10 years, AI electronic medical records via speech recognition will be widespread and considered normal. I have high expectations for the significant role the medical-specialized mocoVoice will play in that.

Onishi: Thank you very much. We too are always aiming for "AI tailored to the frontlines."

Saeki: With doctor overwork and working styles becoming such major issues, I hope these initiatives will ease the burden on the frontlines even a little.

Onishi: Thank you for your valuable insights today.


About the mocoVoice Medical Model

In addition to the standard model equipped in mocoVoice, we have newly trained it on approximately 140,000 words recorded in the Medical Concept and Knowledge Linking Database JMED-DICT mini. Furthermore, we are working to improve accuracy in collaboration with Dr. Saeki, Lecturer/Clinical Associate Professor at Hiroshima University Hospital, allowing you to use industry-leading high-accuracy transcription. We also offer on-premises/on-device provision.

For usage inquiries, please consult with us via our Contact Form.

Related Links

Contact