🔗 원문 전체 보기 → Venturesquare.net

요약
가우디오랩이 두세 달 걸리던 영화·드라마 더빙·자막 작업을 일주일로 단축한 AI 콘텐츠 현지화 플랫폼 'GSP(가우디오 스튜디오 프로)'를 선보였다. 마스터 파일 하나에서 대사·음악·효과음(D/M/E)을 분리해 더빙·자막·음악 교체·영상 편집까지 원스톱으로 처리한다. 자체 음원 분리 엔진 GSEP과 46만 건 규모 저작권 클린 음악 라이브러리, 음량평준화 기술 LM1이 결합됐고, AI가 전 과정을 맡고 사람이 검수를 담당하는 HITL 방식으로 운영된다. 202…
본문
-두세 달 걸리던 더빙·자막 작업을 1주일로…AI 음원 분리 기술이 만든 콘텐츠 현지화의 새 공식
-D/M/E 분리부터 더빙·자막·음악 교체까지 원스톱…마스터 파일 하나로 구작도 글로벌 배급 가능
-2026 대한민국 임팩테크 대상 대통령상·CES 혁신상 2관왕…미국·일본·중국·영국·독일·벨기에 등 글로벌 거점 확보
영화나 드라마를 해외에 수출하려면 더빙과 자막이 필수다. 성우 섭외부터 녹음, 검수, 음악 저작권 처리, 영상 편집까지 보통 두 달이 넘게 걸리는 작업이다. 가우디오랩(Gaudio Lab)은 지난달 이 과정을 일주일로 단축한 AI 콘텐츠 현지화 플랫폼 ‘가우디오 스튜디오 프로(Gaudio Studio Pro, 이하 GSP)’를 선보였다.
가우디오랩은 2015년 헤드폰 공간 음향(바이노럴 렌더링) 기술이 국제표준으로 채택된 것을 계기로 설립된 AI 오디오 기술 스타트업이다. 음향공학 박사 8명을 포함해 40여 명의 오디오 전문가가 함께하고 있다. 네이버, MBC, SBS, TVING, 현대모비스, LG전자, CJ, NHN, 멜론, 라인뮤직 등 국내외 유수 기업이 가우디오랩의 기술로 매일 5,000만 명이 넘는 사용자에게 좋은 소리 경험을 전하고 있다.
기술력은 일찌감치 검증됐다. 2023년부터 2026년까지 CES 혁신상을 4년 연속 수상하며 누적 6관왕에 올랐고, 2024년에는 SXSW 혁신상 파이널리스트에 이름을 올렸다. 앞서 2017년 런던 VR Awards에서 ‘올해의 최고 VR 혁신 기업상’을 받기도 했다. 표준화 영역에서도 성과가 뚜렷하다. 2013년과 2018년 두 차례에 걸쳐 핵심 기술이 ISO/IEC MPEG-H 국제표준으로 채택됐고, 2022년에는 CES를 주최하는 미국 CTA의 ANSI/CTA 표준에도 반영됐다. 한국 오디오 기술이 세계 표준의 한 축으로 자리 잡았다는 평가를 받았다. 2026년에는 대한민국 임팩테크 대상에서 최고 영예인 대통령상을 거머쥐며 기술력과 사업성을 동시에 인정받았다.
가우디오랩의 사업은 세 축으로 나뉜다. 콘텐츠 현지화 플랫폼 ‘GSP’, 원곡 기반 노래방 솔루션 ‘가우디오 씽(Gaudio Sing)’, 그리고 음량평준화·스페이셜 업믹스·바이노럴 렌더링으로 대표되는 기술군이다.
가우디오 씽은 MIDI 반주가 표준이던 노래방 시장에 AI 음원 분리 기술을 적용해 원곡 그대로의 고품질 반주를 구현했다. 핵심 기능은 세 가지다. 마이크 입력에 맞춰 보컬 음량을 자동 조절하는 ‘스마트 필(Smart Fill™)’, 피치·타이밍·음색·비브라토·표현력을 분석하는 정밀 채점 시스템 ‘트루스코어(TrueScore™)’, 글자 단위 가사 동기화 기술 ‘레터 싱크(Letter Sync)’다. 최근에는 일본 시장과 차량용 인포테인먼트 ‘가우디오 씽 포 카’로 영역을 넓히고 있다.
가우디오랩이 개발한 음량평준화 기술은 한국정보통신기술협회(TTA) 국내 표준과 미국 CTA 표준으로 채택돼 네이버의 모든 동영상·오디오 플랫폼은 물론 벅스, 플로 등에서 쓰이고 있다. 스페이셜 업믹스 기술은 네이버 VIBE, LG전자 스마트폰, 베트남 빈스마트 등에서 검증을 마쳤다. OTT와 스트리밍, 자동차, 메타버스, 영화관까지. 소리가 있는 어디에나 가우디오랩의 기술이 닿아 있는 셈이다.
가우디오랩은 AI 음원 분리 엔진 ‘GSEP(Gaudio Source SEParation)’을 활용해 콘텐츠 수출 시 부딪히는 음악 저작권 문제를 풀어주는 ‘뮤직 리플레이스먼트’ 사업을 먼저 시작했고, 그 과정에서 더빙과 자막 수요가 함께 있다는 사실을 확인했다. 지난해 K-FAST 사업에서 100편이 넘는 콘텐츠에 AI 더빙을 적용하며 시장성도 검증했다. 이렇게 다져진 기술과 경험을 한데 묶은 결과물이 GSP다. 정식 론칭에 앞서 지난 4월 라스베이거스에서 열린 ‘NAB Show 2026’에서 처음 공개했고, 이 자리에서만 약 200곳의 고객사와 접점을 만들었다.
가우디오랩의 김병정 PO를 만나 GSP의 특징, 기술, 사업 이야기를 들었다. 김 PO는 컴퓨터사이언스 전공자로 11년간 개발자로 일한 뒤 IoT 플랫폼과 채팅 상담 플랫폼 스타트업을 공동 창업했고, 지난해 8월 가우디오랩에 합류해 GSP의 개발과 론칭을 책임지고 있다.

잠자던 구작도 한 번에…마스터 파일 하나로 끝내는 글로벌 현지화
“음식 배달 앱처럼 주문하면 완제품이 나오는 서비스입니다.”
김 PO는 GSP를 한마디로 이렇게 설명했다. 방송국이나 스튜디오가 자사의 콘텐츠를 GSP에 올리면, 며칠 뒤 더빙과 자막, 음악 교체, 영상 편집까지 마친 현지화 결과물을 받아볼 수 있다.
이 서비스의 핵심은 원본 마스터 파일 하나만으로 작업이 가능하다는 점이다. 최신 콘텐츠는 대사·음악·효과음(Dialogue·Music·Effects, D/M/E) 트랙이 분리된 상태로 보관되지만, 구작은 셋이 하나로 믹싱된 마스터 파일만 남아 있는 경우가 많다. D/M/E가 분리되지 않으면 더빙도, 자막도, 음악 교체도 사실상 불가능하다. GSP는 마스터 파일에서 D/M/E를 깨끗하게 추출한다. 분리된 대사 트랙은 음성-텍스트 변환(STT)을 거쳐 텍스트로 옮겨진 뒤 현지 언어로 번역된다. 단순 번역이 아니라 타임코드가 찍힌 대본으로 가공되는 ‘미디어 프리프로덕션’ 단계다. 이 대본은 더빙의 토대이자 자막의 원본으로 함께 쓰인다. AI는 이 대본을 바탕으로 음성을 합성(TTS)해 1차 더빙본을 만들고, 말의 속도와 캐릭터별 보이스 일관성, 음질 등의 미세한 이슈는 사람이 검수한다. 가우디오랩이 강조하는 HITL(Human-in-the-Loop) 방식이다.
“아직까지 AI만으로는 100% 처리가 어렵습니다. 사람이 중간중간 품질을 검수해야 일관된 결과물이 나옵니다. 그래서 DPQC(더빙 프로덕션 퀄리티 컨트롤)가 반드시 필요합니다.”
음악 처리도 별도 공정이다. 분리된 배경음악은 가우디오랩이 보유한 약 46만 건의 ‘저작권 클린’ 음악 라이브러리에서 같은 무드의 곡으로 자동 교체된다. 사람이 작곡한 ‘휴먼메이드’ 음원으로 구성된 라이브러리여서 저작권 문제를 깔끔하게 풀 수 있다. 영상 후반 작업도 GSP가 함께 맡는다. 더빙된 음성에 맞춰 배우의 입술 모양을 자연스럽게 조정하는 립싱크 기술, 수출 시 노출되면 안 되는 상품 로고나 특정 장면을 가리는 영상 블러·컷 기능까지 갖췄다.
작업이 끝난 뒤에도 사용자는 GSP 자체 플레이어에서 타임라인을 따라가다 마음에 들지 않는 구간에 코멘트를 남길 수 있다. 코멘트는 곧바로 작업자에게 전달돼 수정·보완으로 이어진다. 피드백 루프가 제품에 내장된 구조다. 이 과정을 모두 거치면 보통 두세 달이 걸리던 영화 한 편의 현지화가 일주일로 줄어든다. 처리 시간이 약 90% 단축되고, 인력과 비용도 함께 줄어든다.
‘세계 1위‘ 음원 분리 GSEP, 자동 평가 DQE, 음량평준화 LM1까지
GSP의 기반 기술은 가우디오랩이 독자 개발한 AI 음원 분리 엔진 GSEP다. 가우디오랩은 MusicRadar, MusicTech, LANDR 등 해외 유수 미디어로부터 음원 분리 분야 ‘세계 1위’로 평가받았다. 일반적으로 떠올리는 보컬·반주 분리와, 영상 콘텐츠 더빙에 필요한 D/M/E 분리는 난이도가 전혀 다르다.
“상식적으로는 그냥 대사를 분리해 번역해서 다시 입히면 된다고 생각할 수 있는데, 스튜디오 레벨의 고품질 더빙을 만들려면 ‘D’를 깨끗하게 분리하는 것 자체가 엄청 어려운 기술입니다. 배우들이 조용한 환경에서 또박또박 말하는 게 아니라, 바람 소리, 풀 소리, 군중 소리 온갖 종류의 환경음 속에서 대사를 하거든요. 거기서 대사만 깨끗하게 걸러내야 합니다.”
배경에 음악이 깔리고 그 음악에 보컬까지 섞인 장면은 한층 까다롭다. 범용 AI 모델은 보컬과 대사를 같은 ‘목소리’로 인식해 한 트랙으로 묶어버리는데, 보컬이 섞인 대사 트랙은 더빙 작업에 사실상 쓸 수 없다. GSEP는 음악 속 보컬과 영상 속 대사를 변별하도록 특별히 학습됐고, 이 차이가 GSP 결과물의 품질을 가른다.
가우디오랩은 더빙의 마지막 단계에서 ‘얼마나 잘 더빙됐는지’를 자동으로 평가하는 ‘DQE(Dubbing Quality Evaluator)’ AI를 개발 중이다. 원래 배우의 화자 톤과 얼마나 가까운지, 감정 연기는 살아 있는지, 새로 믹스된 오디오의 음질은 깨끗한지처럼 사람이 일일이 들어 판단하던 항목을 AI가 대신한다.
마지막 단계에는 가우디오랩의 또 다른 강점인 음량평준화(LM1) 기술이 적용된다. 채널 간, 콘텐츠와 광고 간 음량 편차로 시청자가 청력 손상을 입는 것을 막는 기술로, GSP를 통과한 콘텐츠는 어느 플랫폼에 올라가든 안정된 음량으로 송출된다.
미주·남미가 먼저 알아본 GSP
GSP는 출시 직전 의미 있는 상을 받았다. 지난 3월 ‘2026 대한민국 임팩테크 대상’ 대통령상이다. 2018년 이후, 대기업이 아닌 기업이 대통령상을 받은 것은 이번이 처음이다. K-콘텐츠의 글로벌 확산이라는 시대적 흐름과 가우디오랩의 기술이 절묘하게 맞아떨어졌다는 평가다. GSP는 CES 2026에서도 영화 제작·배급(Filmmaking & Distribution)과 엔터프라이즈 테크(Enterprise Tech) 두 분야에서 동시에 혁신상을 받으며 기술력을 거듭 인정받았다.
지난 4월 NAB Show 2026에서 GSP를 글로벌 무대에 처음 공개했을 때, 부스를 찾은 고객 가운데 미주와 남미 비중이 두드러졌다. K-콘텐츠에 대한 관심이 높은 데다 자막을 선호하는 한국과 달리 더빙을 선호하는 남미 시장의 특성이 가우디오랩에는 호재로 작용했다.
“고객들이 정말 많이 찾아왔습니다. 3~5분 분량의 더빙 샘플을 받아보는 PoC(Proof of Concept) 단계에서 만족한 고객이 GSP로 이어지고 있습니다.”
이미 한국과 일본의 주요 방송사가 가우디오랩을 현지화 파트너로 선택해 GSP가 정식 론칭 전부터 가동되고 있었고, 이를 통해 송출된 콘텐츠는 미국·일본·중국·영국·독일·벨기에 등으로 흘러 나갔다. SBS ‘런닝맨’이 대표적이다.
다음 과제는 ‘언어’의 폭을 넓히는 일이다. GSP는 현재 29개국 언어의 AI 더빙을 지원하는데, 더빙 검수를 함께할 언어 전문가 풀을 더 두텁게 확보해야 글로벌 고객을 놓치지 않을 수 있다. 가우디오랩이 왜 더빙에 강한 회사인지 증명할 수 있는 데이터와 레퍼런스를 쌓아 가는 일도 함께 진행한다.

K-콘텐츠는 이미 글로벌 유통망 위에 올라 있다. 그러나 그 확장의 폭은 결국 ‘얼마나 빠르게, 얼마나 합리적인 비용으로, 얼마나 높은 품질로’ 현지화할 수 있느냐에 달려 있다. 두 달이 걸리던 작업을 일주일로 줄이고, 한 시간 분량의 영상을 한 시간 안에 처리한다. AI가 워크플로우의 처음과 끝을 책임지고, 사람은 검수를 맡는다. 가우디오랩이 10년간 쌓아온 AI 오디오 기술을 기반으로 탄생한 GSP가 그 해답이 될 수 있다.
‘소리’를 다루는 회사가 콘텐츠의 국경을 지우는 일에 나섰다. 가우디오랩이 콘텐츠 현지화를 통해 K-콘텐츠를 얼마나 보급할지 관심이 주목된다.
Like this:
Like Loading...Related
Gaudio Lab Changes the Landscape of K-Content Exports with AI Content Localization Platform 'GSP'
– Reducing dubbing and subtitling work from two to three months to one week… A new formula for content localization created by AI audio separation technology
– One-stop service from D/M/E separation to dubbing, subtitles, and music replacement… Global distribution of older titles possible with a single master file
– Winner of the 2026 Korea Impact Tech Awards Presidential Award and CES Innovation Award… Securing global bases in the US, Japan, China, UK, Germany, Belgium, and more
Dubbing and subtitles are essential for exporting movies or dramas overseas. This process, which typically takes over two months from casting voice actors to recording, quality control, music copyright processing, and video editing, was introduced last month by Gaudio Lab, an AI content localization platform called 'Gaudio Studio Pro (GSP),' which shortens this process to one week.
Gaudio Lab is an AI audio technology startup established in 2015 following the adoption of headphone spatial acoustics (binaural rendering) technology as an international standard. It is staffed by over 40 audio experts, including eight Ph.D.s in acoustic engineering. Leading domestic and international companies such as Naver, MBC, SBS, TVING, Hyundai Mobis, LG Electronics, CJ, NHN, Melon, and Line Music are delivering excellent sound experiences to over 50 million users every day using Gaudio Lab's technology.
Its technological prowess was proven early on. It won the CES Innovation Award for four consecutive years from 2023 to 2026, accumulating a total of six awards, and was named a finalist for the SXSW Innovation Award in 2024. Prior to that, it received the "Best VR Innovation Company of the Year" award at the London VR Awards in 2017. Achievements in the field of standardization are also notable. Its core technology was adopted as the ISO/IEC MPEG-H international standard on two occasions, in 2013 and 2018, and was also reflected in the ANSI/CTA standard of the US CTA, the organizer of CES, in 2022. This led to the assessment that Korean audio technology has established itself as a key pillar of global standards. In 2026, it clinched the highest honor, the Presidential Award, at the Korea Impact Tech Awards, receiving recognition for both its technological prowess and business potential.
Gaudio Lab's business is divided into three pillars: the content localization platform 'GSP', the original song-based karaoke solution 'Gaudio Sing', and a group of technologies represented by volume leveling, spatial upmixing, and binaural rendering.
Gaudio Think has implemented high-quality backing tracks that faithfully reproduce the original song by applying AI audio separation technology to the karaoke market, where MIDI backing tracks were the standard. It has three core features: 'Smart Fill™,' which automatically adjusts vocal volume based on microphone input; 'TrueScore™,' a precision scoring system that analyzes pitch, timing, timbre, vibrato, and expressiveness; and 'Letter Sync,' a character-level lyric synchronization technology. Recently, it has been expanding its reach into the Japanese market and the in-car infotainment system 'Gaudio Think for Car.'
The volume leveling technology developed by Gaudio Lab has been adopted as a domestic standard by the Korea Information and Communications Technology Association (TTA) and a US CTA standard, and is being used across all of Naver's video and audio platforms, as well as Bugs and Flo. Spatial Upmix technology has been verified on Naver VIBE, LG Electronics smartphones, and VinSmart in Vietnam. From OTT and streaming to automobiles, the metaverse, and movie theaters, Gaudio Lab's technology has essentially reached wherever there is sound.
Gaudio Lab first launched a "Music Replacement" business utilizing its AI audio separation engine, "GSEP (Gaudio Source SEParation)," to resolve music copyright issues encountered during content export; in the process, the company confirmed that there was a simultaneous demand for dubbing and subtitles. Last year, the company also verified market viability by applying AI dubbing to over 100 pieces of content through the K-FAST project. GSP is the result of combining these solidified technologies and experiences. Prior to its official launch, it was unveiled for the first time at the "NAB Show 2026" held in Las Vegas last April, where it established contact with approximately 200 client companies.
We met with Gaudio Lab’s Product Owner, Kim Byeong-jeong, to hear about the features, technology, and business of GSP. Kim, a computer science major, worked as a developer for 11 years before co-founding startups for IoT platforms and chat consultation platforms. He joined Gaudio Lab last August and is responsible for the development and launch of GSP.

All your dormant classics at once … Global localization finished with a single master file
It is a service where you receive a finished product when you place an order, just like a food delivery app.
PO Kim explained GSP in a single sentence as follows: When a broadcaster or studio uploads its content to GSP, they can receive the localized result a few days later, complete with dubbing, subtitles, music replacement, and video editing.
The core of this service is that it allows work to be performed using only a single original master file. While modern content stores Dialogue, Music, and Effects (D/M/E) tracks separately, older works often retain only a master file where the three are mixed together. Without separating the D/M/E, dubbing, subtitling, and music replacement are virtually impossible. GSP cleanly extracts the D/M/E from the master file. The separated dialogue tracks are converted into text via Speech-to-Text (STT) and then translated into the local language. This is not a simple translation, but a "media pre-production" stage where the content is processed into a time-coded script. This script serves as both the foundation for the dubbing and the source for the subtitles. Based on this script, AI synthesizes speech (TTS) to create a first-round dubbed version, while humans review minute details such as speech speed, voice consistency across characters, and sound quality. This is the Human-in-the-Loop (HITL) method emphasized by Gaudio Lab.
"Currently, it is difficult to process 100% using AI alone. Human quality checks must be performed periodically to ensure consistent results. Therefore, DPQC (Dubbing Production Quality Control) is absolutely necessary."
Music processing is also a separate process. The separated background music is automatically replaced with songs of the same mood from Gaudio Lab's library of approximately 460,000 "copyright-clean" music tracks. Since the library consists of "human-made" audio composed by humans, copyright issues can be resolved cleanly. GSP also handles video post-production. It is equipped with lip-sync technology that naturally adjusts the actor's lip movements to match the dubbed voice, as well as video blur and cut functions that conceal product logos or specific scenes that should not be exposed during export.
Even after the work is finished, users can follow the timeline within the GSP player and leave comments on sections they are not satisfied with. These comments are immediately forwarded to the workers, leading to revisions and improvements. The feedback loop is built into the product. By going through this entire process, the localization of a movie, which typically takes two to three months, is reduced to one week. Processing time is shortened by approximately 90%, and manpower and costs are reduced as well.
' World's No. 1 ' GSEP audio separation , DQE automatic evaluation , and up to LM1 volume leveling
The underlying technology of GSP is GSEP, an AI audio separation engine independently developed by Gaudio Lab. Gaudio Lab has been recognized as the "world's number one" in the field of audio separation by leading international media outlets such as MusicRadar, MusicTech, and LANDR. The difficulty level of vocal and instrumental separation, as generally imagined, is completely different from the D/M/E separation required for video content dubbing.
"Common sense might lead you to think that you just need to separate the dialogue, translate it, and paste it back in, but to create high-quality studio-level dubbing, cleanly separating the 'D' sound itself is an incredibly difficult technique. Actors aren't speaking clearly in a quiet environment; instead, they deliver their lines amidst all kinds of ambient noise—wind, grass, crowds—and so on. You have to cleanly filter out just the dialogue from that environment."
Scenes with background music mixed with vocals are particularly tricky. General-purpose AI models recognize vocals and dialogue as the same 'voice' and group them into a single track, making dialogue tracks mixed with vocals virtually unusable for dubbing. GSEP is specially trained to distinguish between vocals in music and dialogue in video, and this difference determines the quality of the GSEP output.
Gaudio Lab is developing a 'DQE (Dubbing Quality Evaluator)' AI that automatically evaluates 'how well dubbing has been done' during the final stage of dubbing. The AI takes over the task of manually judging criteria that were previously assessed by humans, such as how closely the audio matches the original actor's tone, whether the emotional performance is authentic, and the clarity of the newly mixed audio.
In the final stage, Gaudio Lab's other strength, Volume Leveling (LM1) technology, is applied. This technology prevents viewers from suffering hearing damage due to volume discrepancies between channels and between content and advertisements; as a result, content that passes through GSP is transmitted at a stable volume regardless of the platform it is uploaded to.
GSP, recognized first by the Americas and South America
GSP received a significant award just before its launch: the Presidential Award at the '2026 Korea Impact Tech Awards' last March. This marks the first time since 2018 that a company other than a large conglomerate has received the Presidential Award. It is seen as a perfect alignment between the current trend of the global expansion of K-content and Gaudio Lab's technology. GSP further solidified its technological prowess at CES 2026 by winning Innovation Awards simultaneously in both the Filmmaking & Distribution and Enterprise Tech categories.
When GSP was first unveiled to the global stage at NAB Show 2026 last April, the proportion of visitors from the Americas and South America was notable among the booth's visitors. This was a favorable factor for Gaudio Lab, given the high interest in K-content and the characteristics of the South American market, which prefers dubbing unlike Korea, where subtitles are preferred.
"We have had a great number of customers. Customers who were satisfied during the Proof of Concept (PoC) stage, where they received 3 to 5-minute dubbing samples, are moving on to GSP."
Major broadcasters in Korea and Japan had already selected Gaudio Lab as their localization partner, and GSP was in operation even before its official launch; content transmitted through this channel flowed into the United States, Japan, China, the United Kingdom, Germany, Belgium, and other countries. SBS’s “Running Man” is a prime example.
The next task is to expand the scope of 'languages'. GSP currently supports AI dubbing in 29 languages, but we must secure a deeper pool of language experts to review the dubbing in order to avoid losing global clients. We are also working to build data and references that prove why Gaudio Lab is a company strong in dubbing.

K-content is already established within global distribution networks. However, the scope of its expansion ultimately depends on how quickly, at a reasonable cost, and with what high quality it can be localized. Tasks that used to take two months are reduced to a week, and a one-hour video is processed within an hour. AI takes charge of the workflow from start to finish, while humans handle the quality control. GSP, born from the AI audio technology Gaudio Lab has accumulated over the past 10 years, can be the answer.
A company dealing with 'sound' has set out to erase the borders of content. Attention is focused on how widely Gaudio Lab will distribute K-content through localization.
GAUDIOLAB、AIコンテンツローカライズプラットフォーム「GSP」でKコンテンツ輸出の版を変える
-2~3ヶ月かかるダビング・字幕作業を1週間で… AI音源分離技術が作成したコンテンツローカライゼーションの新しい公式
-D/M/E分離からダビング・字幕・音楽交換までワンストップ…マスターファイル一つで旧作もグローバル配信可能
-2026 韓国インパックテク大統領賞・CES革新賞2冠王…アメリカ・日本・中国・イギリス・ドイツ・ベルギーなどグローバル拠点確保
映画やドラマを海外に輸出するにはダビングと字幕が必須だ。声優交渉から録音、検収、音楽著作権処理、映像編集まで、通常2ヶ月以上かかる作業だ。 GAUDIOLAB(Gaudio Lab)は先月この過程を一週間に短縮したAIコンテンツローカライズプラットフォーム「ガウディオスタジオプロ(Gaudio Studio Pro、以下GSP)」を披露した。
GAUDIOLABは、2015年にヘッドフォン空間音響(バイノーラルレンダリング)技術が国際標準として採用されたことをきっかけに設立されたAIオーディオ技術のスタートアップだ。音響工学博士8人を含めて40人余りのオーディオ専門家が共同している。ネイバー、MBC、SBS、TVING、現代モービス、LG電子、CJ、NHN、メロン、ラインミュージックなど国内外有数企業がGAUDIOLABの技術で毎日5,000万人を超えるユーザーに良い音経験を伝えている。
技術力は早く検証された。 2023年から2026年までCESイノベーション賞を4年連続受賞し、累積6冠王に上がり、2024年にはSXSWイノベーション賞ファイナリストに名を連ねた。これに先立ち、2017年ロンドンVR Awardsで「今年の最高VRイノベーション企業賞」を受けた。標準化領域でも成果がはっきりしている。 2013年と2018年の2回にわたってコア技術がISO/IEC MPEG-H国際標準として採用され、2022年にはCESを主催する米国CTAのANSI/CTA規格にも反映された。韓国オーディオ技術が世界標準の一軸となったという評価を受けた。 2026年には大韓民国インパックテク大賞で最高栄誉人大統領賞を握り、技術力と事業性を同時に認められた。
GAUDIOLABの事業は3軸に分かれています。コンテンツローカライゼーションプラットフォーム「GSP」、原曲ベースのカラオケソリューション「ガウディオ・シン(Gaudio Sing)」、そして音量平準化・スペーシャルアップミックス・バイノーラルレンダリングに代表される技術群だ。
ガウディオ・シンはMIDI伴奏が標準だったカラオケ市場にAI音源分離技術を適用し、原曲そのままの高品質伴奏を実現した。重要な機能は3つあります。マイク入力に合わせてボーカル音量を自動調整する「スマートフィル(Smart Fill™)」、ピッチ・タイミング・音色・ビブラート・表現力を分析する精密採点システム「TrueScore™」、文字単位歌詞同期技術「レターシンク(Letter Sync)」だ。最近では日本市場と車両用インフォテインメント「ガウディオ・シンポカー」で領域を広げている。
GAUDIOLABが開発した音量平準化技術は韓国情報通信技術協会(TTA)国内標準と米国CTA標準として採用され、ネイバーのすべての動画・オーディオプラットフォームはもちろんバックス、フローなどで使われている。スペシャルアップミックス技術はネイバーVIBE、LG電子スマートフォン、ベトナムビンスマートなどで検証を終えた。 OTTとストリーミング、車、メタバス、映画館まで。音があるどこにもGAUDIOLABの技術が届いているわけだ。
GAUDIOLABはAI音源分離エンジン「GSEP(Gaudio Source SEParation)」を活用してコンテンツ輸出時にぶつかる音楽著作権問題を解く「ミュージックリプレイスメント」事業を先に開始し、その過程でダビングと字幕需要が一緒にあることを確認した。昨年、K-FAST事業で100本を超えるコンテンツにAIダビングを適用し、市場性も検証した。このように固められた技術と経験を結んだ結果物がGSPだ。正式ローンチに先立ち、4月にラスベガスで開かれた「NAB Show 2026」で初めて公開し、この場でのみ約200社の顧客と接点を作った。
GAUDIOLABのキム・ビョンジョンPOに会い、GSPの特徴、技術、事業の話を聞いた。キムPOはコンピュータサイエンス専攻者として11年間開発者として働いた後、IoTプラットフォームとチャット相談プラットフォームスタートアップを共同創業し、昨年8月GAUDIOLABに合流してGSPの開発とローンチを担当している。

寝ていた旧作も一度に…マスターファイル1つで終わるグローバルローカライゼーション
「食品配信アプリのように注文すると完成品が出るサービスです。」
キムPOはGSPを一言でこう説明した。放送局やスタジオが自社のコンテンツをGSPに載せれば、数日後にダビングや字幕、音楽の入れ替え、映像編集まで終えたローカライズ結果を受け取ることができる。
このサービスの核心は、元のマスターファイル1つだけで作業が可能だという点だ。最新のコンテンツは大使・音楽・効果音(Dialogue・Music・Effects、D/M/E)トラックが分離された状態で保存されるが、旧作は三つ一つにミキシングされたマスターファイルだけが残っている場合が多い。 D/M/Eが分離されなければダビングも、字幕も、音楽交換も事実上不可能だ。 GSPはマスターファイルからD / M / Eをきれいに抽出します。分離された代謝トラックは、音声テキスト変換(STT)を経てテキストに移動された後、ローカル言語に翻訳される。単純翻訳ではなく、タイムコードが撮られた台本に加工される「メディアプリプロダクション」の段階だ。この台本はダビングの土台であり、字幕の原本として一緒に使われる。 AIはこの台本をもとに音声を合成(TTS)して1次ダビングボーンを作り、馬の速度とキャラクター別ボイス一貫性、音質などの微細な問題は人が検収する。 GAUDIOLABが強調するHITL(Human-in-the-Loop)方式だ。
「まだAIだけでは100%処理が難しいです。人が中中間品質を検収しなければ一貫した結果が出ます。だからDPQC(ダビングプロダクションクオリティコントロール)が必ず必要です。」
音楽処理も別途工程だ。分離されたバックグラウンドミュージックは、GAUDIOLABが保有する約46万件の「著作権クリーン」音楽ライブラリで同じムードの曲に自動的に置き換えられる。人が作曲した「ヒューマンメイド」音源で構成されたライブラリなので著作権問題をきれいに解くことができる。映像後半の作業もGSPが一緒に務める。ダビングされた音声に合わせて俳優の唇の形を自然に調整するリップシンク技術、輸出時に露出してはならない商品のロゴや特定のシーンを隠す映像ブラー・カット機能まで備えた。
作業が終わった後も、ユーザーはGSP自体プレイヤーでタイムラインをたどり、気に入らない区間にコメントを残すことができる。コメントはすぐに作業者に伝えられ、修正・補完につながる。フィードバックループが製品に組み込まれた構造だ。この過程をすべて経ると、普通2~3ヶ月かかった映画一本のローカライゼーションが一週間に減る。処理時間が約90%短縮され、人員とコストも一緒に減る。
「世界1位」音源分離GSEP、自動評価DQE、音量平準化LM1まで
GSPの基盤技術はGAUDIOLABが独自開発したAI音源分離エンジンGSEPだ。 GAUDIOLABはMusicRadar、MusicTech、LANDRなど海外有数メディアから音源分離分野「世界1位」と評価された。一般的に思い浮かぶボーカル・伴奏分離と、映像コンテンツダビングに必要なD/M/E分離は難易度がまったく異なる。
「常識的にはただのセリフを分離して翻訳して再び塗ればいいと思うが、スタジオレベルの高品質ダビングを作るには「D」をきれいに分離すること自体がものすごく難しい技術です。そこでセリフだけをきれいにろ過しなければなりません。」
背景に音楽が敷かれ、その音楽にボーカルまで混ざったシーンは一層トリッキーだ。汎用AIモデルはボーカルとセリフを同じ「声」として認識して1トラックに結びついてしまうが、ボーカルが混ざったセリフトラックはダビング作業に事実上書けない。 GSEPは音楽の中のボーカルと映像の中のセリフを弁別するように特別に学習され、この違いがGSP結果の品質を分ける。
GAUDIOLABはダビングの最後の段階で「どれだけうまくダビングされたか」を自動的に評価する「DQE(Dubbing Quality Evaluator)」AIを開発中だ。もともと俳優の話者のトーンにどれだけ近いのか、感情演技は生きているのか、新しくミックスされたオーディオの音質はきれいなかのように人が毎日聞いて判断していた項目をAIが代わる。
最後のステップでは、GAUDIOLABのもう1つの強みである音量平準化(LM1)技術が適用されます。チャンネル間、コンテンツと広告間の音量偏差で視聴者が聴力損傷を受けるのを防ぐ技術で、GSPを通過したコンテンツはどのプラットフォームに上がっても安定した音量で送出される。
米州・南米が先に調べたGSP
GSPは発売直前に意味のある賞を受賞しました。去る3月'2026大韓民国インパックテク大賞'大統領賞だ。 2018年以降、大企業ではない企業が大統領賞を受けたのは今回が初めてだ。 K-コンテンツのグローバル拡散という時代的流れとGAUDIOLABの技術が絶妙に合致したという評価だ。 GSPはCES 2026でも映画製作・配給(Filmmaking & Distribution)とエンタープライズテック(Enterprise Tech)の2分野で同時に革新賞を受け、技術力を重ね認められた。
去る4月NAB Show 2026でGSPをグローバル舞台に初公開した時、ブースを訪れた顧客の中で米州と南米の比重が目立った。 K-コンテンツへの関心が高いうえ、字幕を好む韓国とは異なり、ダビングを好む南米市場の特性がGAUDIOLABには好材料として作用した。
「顧客が本当にたくさん訪れてきました。3~5分分のダビングサンプルを受け取るPoC(Proof of Concept)段階で満足した顧客がGSPにつながっています。」
すでに韓国と日本の主要放送会社がGAUDIOLABをローカライズパートナーとして選択し、GSPが正式ローンチ前から稼働しており、これを通じて送出されたコンテンツは米国・日本・中国・イギリス・ドイツ・ベルギーなどに流れている。 SBS「ランニングマン」が代表的だ。
次の課題は「言語」の幅を広げることだ。 GSPは現在29カ国の言語のAIダビングをサポートしています。 GAUDIOLABがなぜダビングに強い会社なのかを証明できるデータとリファレンスを積んでいくことも一緒に進む。

K-コンテンツはすでにグローバル流通網の上に上がっている。しかし、その拡張の幅は、結局「どれだけ速く、どれだけ合理的なコストで、どれだけ高品質で」ローカライズできるかによって決まる。 2ヶ月かかる作業を1週間に減らし、1時間分の映像を1時間で処理する。 AIがワークフローの初めと終わりを担当し、人は検収を担当する。 GAUDIOLABが10年間積み重ねてきたAIオーディオ技術を基盤に誕生したGSPがその答えになることができる。
「音」を扱う会社がコンテンツの国境を消去することに乗り出した。 GAUDIOLABがコンテンツのローカライゼーションを通じてK-コンテンツをどれだけ普及するか関心が注目される。
Gaudio Lab 利用人工智能内容本地化平台“GSP”改变韩语内容出口格局
将配音和字幕制作周期从两到三个月缩短到一周……人工智能音频分离技术创造了一种全新的内容本地化方案
– 提供从D/M/E分离到配音、字幕和音乐替换的一站式服务……只需一个母带文件即可实现老片的全球发行
荣获2026年韩国影响力科技奖总统奖和CES创新奖……在美国、日本、中国、英国、德国、比利时等国家建立全球业务基地
配音和字幕对于电影或电视剧出口海外至关重要。通常情况下,从选角、录音、质量控制、音乐版权处理到视频剪辑,整个流程需要两个多月的时间。而上个月,Gaudio Lab推出了一款名为“Gaudio Studio Pro (GSP)”的人工智能内容本地化平台,将这一流程缩短至一周。
Gaudio Lab是一家人工智能音频技术初创公司,成立于2015年,其成立的背景是耳机空间声学(双耳渲染)技术已被国际认可。公司拥有超过40位音频专家,其中包括8位声学工程博士。包括Naver、MBC、SBS、TVING、现代摩比斯、LG电子、CJ、NHN、Melon和Line Music在内的国内外领先企业,每天都在使用Gaudio Lab的技术,为超过5000万用户提供卓越的音频体验。
其技术实力早已得到充分证明。该公司从2023年至2026年连续四年荣获CES创新奖,累计获得六项大奖,并于2024年入围SXSW创新奖决赛。此前,该公司在2017年伦敦VR大奖中荣获“年度最佳VR创新公司”奖。其在标准化领域的成就也同样显著。其核心技术曾两次被采纳为ISO/IEC MPEG-H国际标准(分别在2013年和2018年),并于2022年被CES主办方美国CTA的ANSI/CTA标准所采纳。这使得韩国音频技术被公认为已成为全球标准的重要支柱。2026年,该公司荣获韩国科技影响力奖最高荣誉——总统奖,其技术实力和商业潜力均得到认可。
Gaudio Lab 的业务分为三大支柱:内容本地化平台“GSP”、原创的基于歌曲的卡拉OK解决方案“Gaudio Sing”以及以音量均衡、空间混音和双耳渲染为代表的一组技术。
Gaudio Think 将 AI 音频分离技术应用于卡拉 OK 市场,取代了以往以 MIDI 伴奏为主的模式,实现了能够忠实还原原曲的高品质伴奏。其三大核心功能包括:根据麦克风输入自动调节人声音量的“Smart Fill™”;能够分析音高、节奏、音色、颤音和表现力的精准评分系统“TrueScore™”;以及能够实现字符级歌词同步的“Letter Sync”。近期,Gaudio Think 已将业务拓展至日本市场,并推出了车载信息娱乐系统“Gaudio Think for Car”。
Gaudio Lab 开发的音量均衡技术已被韩国信息通信技术协会 (TTA) 和美国 CTA 采纳为国内标准,并应用于 Naver 的所有视频和音频平台,以及 Bugs 和 Flo 等应用。空间混音技术已在 Naver VIBE、LG 电子智能手机和越南 VinSmart 上得到验证。从 OTT 和流媒体到汽车、元宇宙和电影院,Gaudio Lab 的技术几乎渗透到了所有有声音的地方。
Gaudio Lab 最初利用其人工智能音频分离引擎“GSEP(Gaudio Source SEParation)”推出“音乐替换”业务,旨在解决内容导出过程中遇到的音乐版权问题;在此过程中,该公司确认市场对配音和字幕服务存在同步需求。去年,该公司还通过 K-FAST 项目,将人工智能配音应用于 100 多部作品,验证了该业务的市场可行性。GSP 正是融合了这些成熟的技术和经验的成果。在正式发布之前,GSP 于去年四月在拉斯维加斯举办的“NAB Show 2026”展会上首次亮相,并与约 200 家客户公司建立了联系。
我们采访了 Gaudio Lab 的产品负责人金炳正 (Kim Byeong-jeong),了解了 GSP 的功能、技术和商业模式。金炳正拥有计算机科学学位,在联合创办物联网平台和在线咨询平台初创公司之前,曾担任开发人员长达 11 年。他于去年 8 月加入 Gaudio Lab,负责 GSP 的开发和发布。

所有尘封已久的经典作品一次性上线……只需一个主文件即可完成全球本地化
这是一种下单后即可收到成品的服务,就像外卖应用程序一样。
PO Kim 用一句话解释了 GSP:当广播公司或工作室将其内容上传到 GSP 时,几天后他们就可以收到本地化的结果,包括配音、字幕、音乐替换和视频编辑。
这项服务的核心在于它仅需一个原始母带文件即可完成所有工作。现代内容通常将对话、音乐和音效 (D/M/E) 音轨分开存储,而老作品往往只保留一个母带文件,其中三者混杂在一起。如果不分离 D/M/E,配音、字幕和音乐替换几乎不可能完成。GSP 可以从母带文件中干净利落地提取 D/M/E。分离出的对话音轨通过语音转文本 (STT) 转换为文本,然后翻译成当地语言。这并非简单的翻译,而是一个“媒体预制作”阶段,在此阶段,内容会被处理成带有时间码的脚本。该脚本既是配音的基础,也是字幕的来源。基于此脚本,人工智能 (AI) 会合成语音 (TTS) 以生成初稿配音版本,同时人工会审核语速、角色语音一致性和音质等细节。这就是 Gaudio Lab 所强调的“人机协同 (HITL)”方法。
“目前,仅靠人工智能很难做到100%处理。必须定期进行人工质量检查,以确保结果的一致性。因此,配音制作质量控制(DPQC)是绝对必要的。”
音乐处理也是一个独立的流程。分离出的背景音乐会自动替换为 Gaudio Lab 音乐库中约 46 万首“无版权”曲目中风格相近的歌曲。由于该音乐库中的音频均由人类创作,因此可以轻松解决版权问题。GSP 还负责视频后期制作。它配备了唇形同步技术,能够自然地调整演员的唇部动作以匹配配音,同时还具备视频模糊和剪辑功能,可以隐藏产品标识或导出时不应暴露的特定场景。
即使工作完成后,用户仍然可以在GSP播放器中查看进度,并对不满意的部分发表评论。这些评论会立即转发给工作人员,从而进行修改和改进。反馈机制已内置于产品中。通过这一流程,通常需要两到三个月的电影本地化工作缩短至一周。处理时间缩短了约90%,人力和成本也随之降低。
“世界第一”的GSEP音频分离、 DQE自动评估以及高达LM1的音量均衡
GSP 的底层技术是 GSEP,这是一款由 Gaudio Lab 自主研发的 AI 音频分离引擎。Gaudio Lab 已被 MusicRadar、MusicTech 和 LANDR 等国际知名媒体评为音频分离领域的“世界第一”。人声和乐器分离的难度,与视频内容配音所需的 D/M/E 分离截然不同,前者通常被人们所理解。
“常识可能会让你觉得只需要把对话提取出来,翻译一下,然后再粘贴回去就行了,但要制作出高质量的录音棚级配音,清晰地分离出‘D’音本身却是一项极其困难的技术。演员们说话时并非身处安静的环境;相反,他们是在各种环境噪音——风声、草声、人群声等等——的包围中完成台词的。你必须从这些环境中清晰地过滤掉对话。”
背景音乐与人声混杂的场景处理起来尤其棘手。通用人工智能模型会将人声和对话识别为同一个“声音”,并将它们合并成一个音轨,导致人声与对话混杂的音轨几乎无法用于配音。GSEP 经过专门训练,能够区分音乐中的人声和视频中的对话,而这种区分决定了 GSEP 输出的质量。
Gaudio Lab 正在开发一款名为“DQE(配音质量评估器)”的人工智能,它能够在配音的最后阶段自动评估“配音完成得如何”。该人工智能将接管之前由人工评估的标准,例如音频与原演员语调的匹配程度、情感表达是否真实以及新混音音频的清晰度。
在最后阶段,Gaudio Lab的另一项优势——音量均衡(LM1)技术——得以应用。这项技术可以防止因频道间以及内容与广告间音量差异而导致的听力损伤;因此,无论上传到哪个平台,通过GSP传输的内容都能以稳定的音量呈现。
GSP,首先由美洲和南美洲认可
GSP在正式发布前夕荣获一项重要奖项:去年3月举行的“2026韩国影响力科技奖”总统奖。这是自2018年以来,除大型企业集团外,首次有公司获此殊荣。这被视为K-content全球扩张趋势与Gaudio Lab技术的完美契合。在2026年国际消费电子展(CES 2026)上,GSP同时斩获电影制作与发行和企业技术两大类别的创新奖,进一步巩固了其技术实力。
GSP去年四月在NAB Show 2026上首次亮相全球舞台时,展位参观者中来自美洲和南美洲的比例尤为显著。鉴于南美市场对韩流内容的高涨兴趣以及南美市场偏好配音而非字幕(这与韩国市场截然不同),这对Gaudio Lab来说无疑是一个利好因素。
“我们已经拥有大量客户。在概念验证 (PoC) 阶段,客户收到了 3 到 5 分钟的配音样本,他们对我们的服务感到满意,现在正在转向 GSP。”
韩国和日本的主要电视台早已选定Gaudio Lab作为本地化合作伙伴,GSP甚至在正式上线前就已开始运营;通过该频道传输的内容远销美国、日本、中国、英国、德国、比利时等国家。SBS的《Running Man》就是一个典型的例子。
下一步是拓展“语言”范围。GSP目前支持29种语言的AI配音,但我们必须确保拥有更庞大的语言专家团队来审核配音,以避免失去全球客户。我们也在努力积累数据和案例,以证明Gaudio Lab在配音领域的实力。

韩流内容已在全球分发网络中站稳脚跟。然而,其扩展规模最终取决于本地化的速度、成本和质量。过去需要两个月才能完成的任务现在只需一周,一小时的视频处理时间也缩短至一小时。人工智能负责从头到尾的工作流程,而人工则负责质量控制。GSP,源自Gaudio Lab过去十年积累的人工智能音频技术,或许正是解决这一问题的关键。
一家专注于“声音”的公司致力于打破内容的边界。人们关注的焦点在于,Gaudio Lab 将如何通过本地化广泛传播韩流内容。
Gaudio Lab révolutionne l'exportation de contenu coréen grâce à sa plateforme de localisation de contenu par IA « GSP ».
Réduire le temps de doublage et de sous-titrage de deux à trois mois à une semaine… Une nouvelle formule pour la localisation de contenu créée par la technologie de séparation audio par IA
Service complet, de la séparation D/M/E au doublage, aux sous-titres et au remplacement de la musique… Distribution mondiale des titres anciens possible grâce à un seul fichier master.
Lauréat du Prix présidentiel des Korea Impact Tech Awards 2026 et du Prix de l'innovation du CES… Présence internationale établie aux États-Unis, au Japon, en Chine, au Royaume-Uni, en Allemagne, en Belgique et ailleurs.
Le doublage et le sous-titrage sont indispensables pour exporter des films ou des séries à l'international. Ce processus, qui prend généralement plus de deux mois (du casting des comédiens de doublage à l'enregistrement, en passant par le contrôle qualité, le traitement des droits musicaux et le montage vidéo), a été simplifié le mois dernier grâce à Gaudio Lab, qui a introduit une plateforme de localisation de contenu basée sur l'IA appelée « Gaudio Studio Pro (GSP) », réduisant ainsi sa durée à une semaine.
Gaudio Lab est une startup spécialisée dans les technologies audio basées sur l'intelligence artificielle, fondée en 2015 suite à l'adoption de la technologie d'acoustique spatiale pour casques (rendu binaural) comme norme internationale. Elle emploie plus de 40 experts audio, dont huit docteurs en ingénierie acoustique. Des entreprises nationales et internationales de premier plan, telles que Naver, MBC, SBS, TVING, Hyundai Mobis, LG Electronics, CJ, NHN, Melon et Line Music, offrent chaque jour une expérience sonore exceptionnelle à plus de 50 millions d'utilisateurs grâce à la technologie de Gaudio Lab.
Son expertise technologique s'est rapidement affirmée. L'entreprise a remporté le prix de l'innovation du CES quatre années consécutives, de 2023 à 2026, accumulant ainsi six récompenses au total, et a été finaliste du prix de l'innovation du SXSW en 2024. Auparavant, elle avait reçu le prix de la « Meilleure entreprise innovante en réalité virtuelle de l'année » lors des London VR Awards en 2017. Ses réalisations en matière de normalisation sont également remarquables. Sa technologie de base a été adoptée comme norme internationale ISO/IEC MPEG-H à deux reprises, en 2013 et 2018, et a également été intégrée à la norme ANSI/CTA de la CTA américaine, organisatrice du CES, en 2022. Ces succès ont permis de considérer la technologie audio coréenne comme un pilier essentiel des normes mondiales. En 2026, elle a décroché la plus haute distinction, le Prix présidentiel, lors des Korea Impact Tech Awards, récompensant à la fois son expertise technologique et son potentiel commercial.
L'activité de Gaudio Lab repose sur trois piliers : la plateforme de localisation de contenu « GSP », la solution de karaoké originale basée sur des chansons « Gaudio Sing », et un ensemble de technologies représentées par le nivellement du volume, le mixage spatial et le rendu binaural.
Gaudio Think a révolutionné le marché du karaoké, jusqu'alors dominé par les pistes d'accompagnement MIDI, en proposant des pistes d'accompagnement de haute qualité reproduisant fidèlement la chanson originale grâce à une technologie de séparation audio basée sur l'IA. Ses trois fonctionnalités principales sont : « Smart Fill™ », qui ajuste automatiquement le volume vocal en fonction du microphone ; « TrueScore™ », un système de notation précis analysant la hauteur, le rythme, le timbre, le vibrato et l'expressivité ; et « Letter Sync », une technologie de synchronisation des paroles au niveau des caractères. Récemment, la société a étendu sa présence au marché japonais avec le système d'infodivertissement embarqué « Gaudio Think for Car ».
La technologie d'égalisation du volume développée par Gaudio Lab a été adoptée comme norme en Corée par la Korea Information and Communications Technology Association (TTA) et comme norme aux États-Unis par la CTA. Elle est utilisée sur toutes les plateformes vidéo et audio de Naver, ainsi que sur Bugs et Flo. La technologie Spatial Upmix a été testée sur Naver VIBE, les smartphones LG Electronics et VinSmart au Vietnam. Des plateformes OTT et de streaming aux automobiles, en passant par le métavers et les salles de cinéma, la technologie de Gaudio Lab est désormais présente partout où le son est diffusé.
Gaudio Lab a d'abord lancé une activité de « remplacement musical » utilisant son moteur de séparation audio par IA, « GSEP (Gaudio Source SEparation) », afin de résoudre les problèmes de droits d'auteur musicaux rencontrés lors de l'exportation de contenu. L'entreprise a ainsi constaté une demande simultanée pour le doublage et le sous-titrage. L'année dernière, elle a également validé la viabilité du marché en appliquant le doublage par IA à plus de 100 contenus dans le cadre du projet K-FAST. GSP est le fruit de la combinaison de ces technologies et expériences éprouvées. Avant son lancement officiel, la solution a été présentée pour la première fois au « NAB Show 2026 » de Las Vegas en avril dernier, où elle a établi des contacts avec près de 200 entreprises clientes.
Nous avons rencontré Kim Byeong-jeong, chef de produit chez Gaudio Lab, pour en savoir plus sur les fonctionnalités, la technologie et le modèle économique de GSP. Diplômé en informatique, Kim a travaillé comme développeur pendant 11 ans avant de cofonder des startups spécialisées dans les plateformes IoT et les plateformes de consultation par messagerie instantanée. Il a rejoint Gaudio Lab en août dernier et est responsable du développement et du lancement de GSP.

Tous vos classiques oubliés réunis … Localisation mondiale finalisée grâce à un seul fichier maître
Il s'agit d'un service où vous recevez un produit fini lorsque vous passez commande, comme une application de livraison de repas.
PO Kim a expliqué le GSP en une seule phrase : lorsqu’un diffuseur ou un studio télécharge son contenu sur le GSP, il peut recevoir le résultat localisé quelques jours plus tard, avec doublage, sous-titres, remplacement de la musique et montage vidéo.
Ce service repose essentiellement sur sa capacité à réaliser des travaux à partir d'un seul fichier master original. Alors que les contenus modernes stockent les pistes de dialogue, de musique et d'effets (D/M/E) séparément, les œuvres plus anciennes ne conservent souvent qu'un seul fichier master où ces trois éléments sont mixés. Sans cette séparation, le doublage, le sous-titrage et le remplacement de la musique sont quasiment impossibles. GSP extrait proprement les pistes D/M/E du fichier master. Les pistes de dialogue séparées sont converties en texte grâce à la synthèse vocale (STT), puis traduites dans la langue locale. Il ne s'agit pas d'une simple traduction, mais d'une véritable étape de préproduction où le contenu est transformé en un script horodaté. Ce script sert à la fois de base au doublage et de source pour les sous-titres. À partir de ce script, une IA synthétise la parole (TTS) pour créer une première version doublée, tandis que des humains vérifient des détails précis tels que le débit de parole, la cohérence des voix entre les personnages et la qualité sonore. C'est la méthode d'intervention humaine (HITL) préconisée par Gaudio Lab.
« Actuellement, il est difficile de traiter 100 % du doublage uniquement par l'IA. Des contrôles qualité humains doivent être effectués périodiquement pour garantir des résultats cohérents. Par conséquent, le contrôle qualité de la production du doublage (DPQC) est absolument nécessaire. »
Le traitement musical est également un processus distinct. La musique de fond est automatiquement remplacée par des morceaux d'ambiance similaire issus de la bibliothèque Gaudio Lab, qui compte environ 460 000 titres libres de droits. Cette bibliothèque étant composée de morceaux créés par des humains, les problèmes de droits d'auteur sont résolus sans difficulté. GSP gère également la post-production vidéo. Il intègre une technologie de synchronisation labiale qui ajuste naturellement les mouvements des lèvres de l'acteur à la voix doublée, ainsi que des fonctions de floutage et de découpage vidéo permettant de masquer les logos ou certaines scènes à ne pas afficher lors de l'exportation.
Même une fois le travail terminé, les utilisateurs peuvent suivre l'avancement du projet dans le lecteur GSP et laisser des commentaires sur les passages qui ne les satisfont pas. Ces commentaires sont immédiatement transmis aux équipes, ce qui permet d'apporter des corrections et des améliorations. Ce système de retour d'information est intégré au produit. Grâce à ce processus, la localisation d'un film, qui prend généralement deux à trois mois, est réduite à une semaine. Le temps de traitement est raccourci d'environ 90 %, et les besoins en personnel et les coûts sont également réduits.
Séparation audio GSEP « n° 1 mondial » , évaluation automatique DQE et nivellement du volume jusqu’à LM1
La technologie sous-jacente de GSP est GSEP, un moteur de séparation audio basé sur l'IA et développé indépendamment par Gaudio Lab. Gaudio Lab est reconnu comme le leader mondial de la séparation audio par des médias internationaux de premier plan tels que MusicRadar, MusicTech et LANDR. La difficulté de la séparation des voix et des instruments, telle qu'on l'imagine généralement, est totalement différente de la séparation D/M/E requise pour le doublage de contenus vidéo.
On pourrait penser qu'il suffit de séparer les dialogues, de les traduire et de les réintégrer, mais pour obtenir un doublage de qualité professionnelle, isoler proprement le son « D » est une technique extrêmement complexe. Les acteurs ne parlent pas dans un environnement silencieux ; ils débitent leurs répliques au milieu de toutes sortes de bruits ambiants : vent, herbe, foule, etc. Il faut donc filtrer parfaitement les dialogues pour les isoler de ce bruit de fond.
Les scènes mêlant musique de fond et voix sont particulièrement complexes. Les modèles d'IA généralistes considèrent la voix et les dialogues comme une seule et même « voix » et les regroupent en une seule piste, rendant ainsi les pistes de dialogues mélangées à la voix pratiquement inutilisables pour le doublage. GSEP est spécialement entraîné pour distinguer la voix dans la musique et les dialogues dans la vidéo, et cette distinction détermine la qualité du rendu GSEP.
Gaudio Lab développe une IA, « DQE (Évaluateur de Qualité de Doublage) », qui évalue automatiquement la qualité du doublage lors de la phase finale. Cette IA prend en charge l'évaluation manuelle de critères auparavant effectués par des humains, tels que la fidélité de la voix à celle de l'acteur original, l'authenticité de l'interprétation émotionnelle et la clarté du mixage audio.
Dans la dernière étape, l'autre atout de Gaudio Lab, la technologie de nivellement du volume (LM1), est mise en œuvre. Cette technologie protège les téléspectateurs des dommages auditifs dus aux différences de volume entre les chaînes et entre les contenus et les publicités ; ainsi, les contenus transitant par GSP sont diffusés à un volume stable, quelle que soit la plateforme de diffusion.
Le SPG, reconnu en premier lieu par les Amériques et l'Amérique du Sud
GSP a reçu une récompense prestigieuse juste avant son lancement : le Prix Présidentiel lors des « Korea Impact Tech Awards 2026 » en mars dernier. C’est la première fois depuis 2018 qu’une entreprise autre qu’un grand conglomérat reçoit ce prix. Cette distinction illustre parfaitement l’adéquation entre la tendance actuelle à l’expansion mondiale des contenus coréens et la technologie de Gaudio Lab. GSP a par ailleurs confirmé son expertise technologique au CES 2026 en remportant simultanément le Prix de l’Innovation dans les catégories Production et Distribution cinématographiques et Technologies d’Entreprise.
Lors de la première présentation de GSP au NAB Show 2026 en avril dernier, la forte présence de visiteurs d'Amérique du Nord et du Sud sur le stand a été un atout majeur pour Gaudio Lab. Ce fut un facteur favorable, compte tenu du vif intérêt pour les contenus coréens et des spécificités du marché sud-américain, qui privilégie le doublage contrairement à la Corée où les sous-titres sont préférés.
« Nous avons eu un grand nombre de clients. Les clients satisfaits lors de la phase de validation de concept (PoC), où ils ont reçu des échantillons de doublage de 3 à 5 minutes, passent maintenant à GSP. »
Les principales chaînes de télévision coréennes et japonaises avaient déjà choisi Gaudio Lab comme partenaire de localisation, et GSP était opérationnel avant même son lancement officiel ; les contenus diffusés via ce canal étaient accessibles aux États-Unis, au Japon, en Chine, au Royaume-Uni, en Allemagne, en Belgique et dans d’autres pays. L’émission « Running Man » de SBS en est un parfait exemple.
La prochaine étape consiste à étendre la gamme de langues prises en charge. GSP propose actuellement le doublage par IA dans 29 langues, mais nous devons constituer un plus large réseau d'experts linguistiques pour superviser les doublages et ainsi éviter de perdre des clients internationaux. Nous travaillons également à la constitution de données et de références démontrant l'expertise de Gaudio Lab en matière de doublage.

Le contenu coréen est déjà bien implanté au sein des réseaux de distribution mondiaux. Toutefois, son expansion dépendra en fin de compte de sa localisation : rapide, économique et de haute qualité. Des tâches qui prenaient auparavant deux mois sont désormais réduites à une semaine, et une vidéo d'une heure est traitée en une heure. L'IA gère l'ensemble du flux de travail, tandis que l'humain assure le contrôle qualité. GSP, fruit de l'expertise audio en IA développée par Gaudio Lab au cours des dix dernières années, pourrait bien être la solution.
Une entreprise spécialisée dans le son s'est donné pour mission d'abolir les frontières du contenu. L'attention se porte désormais sur l'ampleur de la diffusion des contenus coréens par Gaudio Lab grâce à la localisation.