Sitemap
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Pages
Posts
portfolio
publications
How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics
Published in Interspeech 2023, 2023
Recommended citation: Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki, Detai Xin, Hiroshi Saruwatari. (2023). "How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics." Interspeech 2023, pp. 1085-1089.
Do Learned Speech Symbols Follow Zipf’s Law?
Published in ICASSP 2024, 2024
Recommended citation: Shinnosuke Takamichi, Hiroki Maeda, Joonyong Park, Daisuke Saito, Hiroshi Saruwatari. (2024). "Do Learned Speech Symbols Follow Zipf's Law?" ICASSP 2024, pp. 12526-12530.
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
Published in ISCA INTERSPEECH 2024, 2024
Recommended citation: Kentaro Onda, Joonyong Park, Nobuaki Minematsu, Daisuke Saito. (2024). "A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora." Proc. ISCA INTERSPEECH, 2024.
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
Published in APSIPA ASC 2024, 2024
Recommended citation: Joonyong Park, Daisuke Saito, Nobuaki Minematsu. (2024). "Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model." APSIPA ASC 2024.
Download Paper
Analyzing the Language of Visual Tokens
Published in Submitted to ACL ARR 2025, 2025
Recommended citation: David M. Chan, Rodolfo Corona, Joonyong Park, Cheol Jun Cho, Yutong Bai, Trevor Darrell. (2025). "Analyzing the Language of Visual Tokens." Submitted to ACL ARR, 2025.
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
Published in ICSA SSW 2025, 2025
Recommended citation: Joonyong Park, Kenichi Nakamura. (2025). "EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens." Proc. ICSA SSW, 2025.
Download Paper
Analysing the Language of Neural Audio Codecs
Published in IEEE ASRU 2025, 2025
Recommended citation: Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito, Hiroshi Saruwatari. (2025). "Analysing the Language of Neural Audio Codecs." Proc. IEEE ASRU, 2025.
Download Paper
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script Texts using Speech Self-Supervised Learning and Language Model
Published in APSIPA 2025, 2025
Recommended citation: Joonyong Park, Daisuke Saito, Nobuaki Minematsu. (2025). "MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script Texts using Speech Self-Supervised Learning and Language Model." APSIPA 2025.
Download Paper
