Heeseung Kim

Assistant Professor · University of Seoul

Heeseung Kim

Building more human-like, real-time voice and omni-modal agents.

Portrait of Heeseung Kim

I am an Assistant Professor in the Department of Artificial Intelligence at the University of Seoul (since March 2026), where I lead the Speech & Interactive AI Lab (SIA Lab). My research centers on interactive and proactive AI: real-time, full-duplex voice and omni-modal agents that listen, see, and step in at the right moment, even before being asked. I also work on generative models and on speech, audio, and music.

Previously, I was a Senior Research Engineer at Qualcomm AI Research Korea (until February 2026), where I worked on training full-duplex speech LLMs. I received my Ph.D. from the Data Science & AI Lab (DSAIL) at Seoul National University (SNU), advised by Sungroh Yoon, where my research focused on speech synthesis (text-to-speech, voice conversion) and speech large language and dialog models (speech LLMs), alongside explorations of generative models in other domains (vision, NLP). I received my B.S. in Electrical and Computer Engineering from Seoul National University.

Research Interests

Three threads toward AI that listens, speaks, and interacts like humans.

Illustration for Interactive & Proactive AI

Interactive & Proactive AI

AI that listens, sees, and talks in real time, and knows when to step in before being asked.

  • Full-duplex dialogue
  • Proactive assistance
  • Omni-modal agents
  • Tool use
Illustration for Generative Models

Generative Models

Generative models that create images, video, text, and audio, spanning diffusion, flow, and language models.

  • Image & video
  • Text
  • Audio
  • Diffusion & flow
Illustration for Speech, Audio & Music

Speech, Audio & Music

Generating and understanding speech, audio, and music: text-to-speech, voice personalization, speech LLMs, and music generation.

  • Text-to-speech
  • Speech LLMs
  • Audio understanding
  • Music generation

Publications

* first author (co-first authors share the mark)  ·  † corresponding author

Selected Publications 5

All Publications 29

Speech & Multimodal LLMs 10

Speech 11

Natural Language Processing 3

Computer Vision 3

Others 2

Education

Ph.D. in Electrical and Computer Engineering

Seoul National University · Mar 2019 - Aug 2025

Integrated M.S./Ph.D. program, Data Science & AI Lab (DSAIL), advised by Prof. Sungroh Yoon.

B.S. in Electrical and Computer Engineering

Seoul National University · Mar 2015 - Feb 2019

Graduated Cum Laude.

Awards, Talks & Services

Honors & Awards

Lectures

  • Guest Lecture, Special Topics on Speech Signal Processing, Yonsei University, October 2026
  • Samsung AI Expert Program, Samsung Electronics, April to May 2026 (3-day lecture series: speech processing, speech recognition, and speech synthesis)
  • Guest Lecture, SW Leadership and Entrepreneurship Seminar, Department of Computer Science & Engineering, Ewha Womans University, March 2026

Invited Talks

  • "AI That Helps Before You Ask: Real-Time Proactive AI, Today and Tomorrow", UOS Research Colloquium, University of Seoul, October 2026
  • "Hey Siri: The Present and Future of Voice Assistants", AGILE (Artificial General Intelligence Laboratory), Ewha Womans University, January 2026
  • "Speech Synthesis to Voice Assistant", Supertone, 2025
  • "Latest Trends in Spoken Dialog Models and Voice Agents", Qualcomm, 2024
  • "Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation", HMG TECH SUMMIT, 2024
  • "Speech and Spoken Dialog Modeling", Neosapience, 2024
  • "A Case Study of Research and Development at Seoul National University Using Amazon Mechanical Turk", AWS Summit Seoul, 2024
  • "Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance", Kakao Enterprise, 2022

Academic Services