Martin Lenglet, Olivier Perrotin, Gérard Bailly. A closer look at internal representations of end-to-end Text-to-Speech models: How is phonetic and acoustic information encoded?. Computer Speech and Language, 2026, 100 (October), pp.101985. ⟨10.1016/j.csl.2026.101985⟩. ⟨hal-05572190⟩
Joan Birulés, Alexandre Duroyal, Anne Vilain, Gérard Bailly, Mathilde Fort. French Speakers Prefer Prosody Over Statistics to Segment Speech. Language and Speech, 2025, ⟨10.1177/00238309251374295⟩. ⟨hal-05501575⟩
Léa Haefflinger, Frédéric Elisei, Gérard Bailly. Data-Driven Control of Eye and Head movements for Triadic Human-Robot Interactions. International Journal of Social Robotics, 2025, 17, pp.1075-1096. ⟨10.1007/s12369-025-01245-2⟩. ⟨hal-05046833⟩
Olivier Perrotin, Brooke Stephenson, Silvain Gerber, Gérard Bailly, Simon King. Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023. Computer Speech and Language, 2025, 90 (March), pp.101747. ⟨10.1016/j.csl.2024.101747⟩. ⟨hal-04813569⟩
Hippolyte Fournier, Sina Alisamir, Safaa Azzakhnini, Isabella Zsoldos, Eléonore Trân, et al.. THERADIA WoZ: An Ecological Corpus for Appraisal-based Affect Research in Healthcare. IEEE Transactions on Affective Computing, 2025, 16 (3), pp.2233-2244. ⟨10.1109/TAFFC.2025.3557465⟩. ⟨hal-05046330⟩
Isabella Zsoldos, Eléonore Trân, Hippolyte Fournier, Franck Tarpin-Bernard, Joan Fruitet, et al.. The Value of a Virtual Assistant to Improve Engagement in Computerized Cognitive Training at Home: Exploratory Study. JMIR Rehabilitation and Assistive Technologies, 2024, 11, pp.e48129. ⟨10.2196/48129⟩. ⟨hal-04625198⟩
Gérard Bailly, Erika Godde, Anne-Laure Piat-Marchand, Marie-Line Bosse. Automatic assessment of oral readings of young pupils. Speech Communication, 2022, 138, pp.67-79. ⟨10.1016/j.specom.2022.01.008⟩. ⟨hal-03585934⟩
Erika Godde, Marie-Line Bosse, Gérard Bailly. Pausing and Breathing while Reading Aloud : Development from 2nd to 7th grade. Reading and Writing : an interdisciplinary journal, 2022, 35, pp.1-27. ⟨10.1007/s11145-021-10168-z⟩. ⟨hal-02954071⟩
Erika Godde, Marie-Line Bosse, Gérard Bailly. Échelle Multi-Dimensionnelle de Fluence : nouvel outil d’évaluation de la fluence en lecture prenant en compte la prosodie, étalonné du CE1 à la 5ème. L’année Psychologique/ Trends in Cognitive Psychology, 2021, 121 (2), pp.19-43. ⟨10.3917/anpsy1.212.0019⟩. ⟨hal-02954060⟩
Erika Godde, Marie-Line Bosse, Gérard Bailly. A review of reading prosody acquisition and development. Reading and Writing : an interdisciplinary journal, 2020, 33 (2), pp.399-426. ⟨10.1007/s11145-019-09968-1⟩. ⟨hal-02185896⟩
Jeesun Kim, Gérard Bailly, Chris Davis. Introduction to the special issue on auditory-visual expressive speech and gesture in humans and machines. Speech Communication, 2018, 98, pp.63-67. ⟨10.1016/j.specom.2018.02.001⟩. ⟨hal-01821001⟩
Emilie Gerbier, Gérard Bailly, Marie-Line Bosse. Audio-visual synchronization in reading while listening to texts: Effects on visual behavior and verbal learning. Computer Speech and Language, 2018, 47 (january), pp.79-92. ⟨10.1016/j.csl.2017.07.003⟩. ⟨hal-01575227⟩
Adela Barbulescu, Rémi Ronfard, Gérard Bailly. Which prosodic features contribute to the recognition of dramatic attitudes?. Speech Communication, 2017, 95, pp.78-86. ⟨10.1016/j.specom.2017.07.003⟩. ⟨hal-01643330⟩
Duc-Canh Nguyen, Gérard Bailly, Frédéric Elisei. Pattern Recognition Letters Learning Off-line vs. On-line Models of Interactive Multimodal Behaviors with Recurrent Neural Networks. Pattern Recognition Letters, 2017, 100, pp.29-36. ⟨10.1016/j.patrec.2017.09.033⟩. ⟨hal-01609535⟩
Adela Barbulescu, Rémi Ronfard, Gérard Bailly. A Generative Audio-Visual Prosodic Model for Virtual Actors. IEEE Computer Graphics and Applications, 2017, 37 (6), pp.40-51. ⟨10.1109/MCG.2017.4031070⟩. ⟨hal-01643334⟩
Alaeddine Mihoub, Gérard Bailly, Christian Wolf, Frédéric Elisei. Graphical models for social behavior modeling in face-to face interaction. Pattern Recognition Letters, 2016, 74, pp.82-89. ⟨10.1016/j.patrec.2016.02.005⟩. ⟨hal-01279427⟩
Gérard Bailly. Critical review of the book » Gaze in Human-Robot Communication « . Journal on Multimodal User Interfaces, 2016, 11 (1), pp.113-114. ⟨10.1007/s12193-016-0219-6⟩. ⟨hal-01310308⟩
Thomas Hueber, Gérard Bailly. Statistical Conversion of Silent Articulation into Audible Speech using Full-Covariance HMM. Computer Speech and Language, 2016, 36, pp.274-293. ⟨10.1016/j.csl.2015.03.005⟩. ⟨hal-01228885⟩
Thomas Hueber, Laurent Girin, Xavier Alameda-Pineda, Gérard Bailly. Speaker-Adaptive Acoustic-Articulatory Inversion using Cascaded Gaussian Mixture Regression. IEEE/ACM Transactions on Audio, Speech and Language Processing, 2015, 23 (12), pp.2246-2259. ⟨10.1109/TASLP.2015.2464702⟩. ⟨hal-01231197⟩
Alaeddine Mihoub, Gérard Bailly, Christian Wolf, Frédéric Elisei. Learning multimodal behavioral models for face-to-face social interaction. Journal on Multimodal User Interfaces, 2015, 9 (3), pp.195-210. ⟨10.1007/s12193-015-0190-7⟩. ⟨hal-01170991⟩
Alberto Parmiggiani, Marco Randazzo, Marco Maggiali, Giorgio Metta, Frédéric Elisei, et al.. Design and Validation of a Talking Face for the iCub. International Journal of Humanoid Robotics, 2015, 12 (3), pp.15500026. ⟨10.1142/S0219843615500267⟩. ⟨hal-01228886⟩
Jean-David Boucher, Ugo Pattacini, Amélie Lelong, Gérard Bailly, Frédéric Elisei, et al.. I reach faster when I see you look: Gaze effects in human-human and human-robot face-to-face cooperation. Frontiers in Neurorobotics, 2012, 6 (3), pp.1-11. ⟨10.3389/fnbot.2012.00003⟩. ⟨hal-00694314⟩
Pierre Badin, Christophe Savariaux, Gérard Bailly, Frédéric Elisei, Louis-Jean Boë. Caractérisation des mécanismes de production de la parole: une approche biométrique et modélisatrice mono-locuteur et multi-dispositifs. Biométrie Humaine et Anthropologie – revue de la Société de biométrie humaine, 2012, 30 (1-2), pp.67-77. ⟨hal-00779188⟩
Panikos Heracleous, Pierre Badin, Gérard Bailly, Norihiro Hagita. A pilot study on augmented speech communication based on Electro-Magnetic Articulography. Pattern Recognition Letters, 2011, 32 (8), pp.1119-1125. ⟨10.1016/j.patrec.2011.02.009⟩. ⟨hal-00582589⟩
Gérard Bailly, Stephan Raidt, Frédéric Elisei. Gaze, conversational agents and face-to-face communication. Speech Communication, 2010, 52 (6), pp.598-612. ⟨10.1016/j.specom.2010.02.015⟩. ⟨hal-00480335⟩
Marion Dohen, Jean-Luc Schwartz, Gérard Bailly. Speech and face-to-face communication – An introduction. Speech Communication, 2010, 52 (6), pp.477-480. ⟨10.1016/j.specom.2010.02.016⟩. ⟨hal-00480343⟩
Pierre Badin, Yuliya Tarabalka, Frédéric Elisei, Gérard Bailly. Can you ‘read’ tongue movements ? Evaluation of the contribution of tongue display to speech understanding. Speech Communication, 2010, 52 (6), pp.493-503. ⟨10.1016/j.specom.2010.03.002⟩. ⟨hal-00463621⟩
Sascha Fagel, Gérard Bailly, Barry-John Theobald. Animating virtual speakers or singers from audio: lip-synching facial animation. EURASIP Journal on Audio, Speech, and Music Processing, 2009, pp.ID 826091. ⟨10.1155/2009/826091⟩. ⟨hal-00473027⟩
Gérard Bailly, Oxana Govokhina, Frédéric Elisei, Gaspard Breton. Lip-synching using speaker-specific articulation, shape and appearance models. EURASIP Journal on Audio, Speech, and Music Processing, 2009, Special issue on animating virtual speakers or singers from audio: Lip-synching facial animation, pp.ID 769494. ⟨10.1155/2009/769494⟩. ⟨hal-00447061⟩
Panikos Heracleous, Denis Beautemps, Viet-Anh Tran, Hélène Loevenbruck, Gérard Bailly. Exploiting visual information for NAM recognition. IEICE Electronics Express, 2009, 6 (2), pp.77-82. ⟨10.1587/elex.6.77⟩. ⟨hal-00357985⟩
Gérard Bailly, Frédéric Elisei, Stephan Raidt. Boucles de perception-action et interaction face-à-face. Revue Française de Linguistique Appliquée, 2008, XIII (2), pp.121-131. ⟨hal-00343411⟩
Pierre Badin, Frédéric Elisei, Gérard Bailly, Christophe Savariaux, Antoine Serrurier, et al.. Têtes parlantes audiovisuelles virtuelles : données et modèles articulatoires – applications. Revue de Laryngologie Otologie Rhinologie, 2007, 128 (5), pp.289-295. ⟨hal-00260326⟩
Alice Caplier, Sébastien Stillittano, Oya Aran, Lale Akarun, Gérard Bailly, et al.. Image and Video for hearing impaired people. EURASIP Journal on Image and Video Processing, 2007, 2007, pp.ID 45641. ⟨10.1155/2007/45641⟩. ⟨hal-00275029⟩
Gérard Bailly, Virginie Attina, Cléo Baras, Patrick Bas, Séverine Baudry, et al.. ARTUS: synthesis and audiovisual watermarking of the movements of a virtual agent interpreting subtitling using cued speech for deaf televiewers. Modelling, measurement and control C, 2006, 67SH (2, supplement : handicap), pp.177-187. ⟨hal-00157826⟩
Maxime Berar, Michel Desvignes, Gérard Bailly, Yohan Payan. 3D Semi-Landmark based Statistical Face Reconstruction. Journal of Computing and Information Technology, 2006, 14 (1), pp.31-43. ⟨hal-00080375⟩
Guillaume Gibert, Gérard Bailly, Denis Beautemps, Frédéric Elisei. Analysis and synthesis of the 3D movements of the head, face and hand of a speaker using cued speech. Journal of the Acoustical Society of America, 2005, 118 (2), pp.1144-1153. ⟨hal-00143622⟩
Maxime Berar, Michel Desvignes, Gérard Bailly, Yohan Payan. Missing data estimation using polynomial kernels. Lecture Notes in Computer Science, 2005, 3686, pp.390-399. ⟨hal-00081903⟩
Maxime Berar, Michel Desvignes, Gérard Bailly, Yohan Payan. 3D Meshes Registration : Application to statistical skull model. Lecture Notes in Computer Science, 2004, 3212, pp.100-107. ⟨hal-00081915⟩
Pierre Badin, Gérard Bailly, Lionel Reveret, Monica Baciu, Christoph Segebarth, et al.. Three-dimensional linear articulatory modeling of tongue, lips and face, based on MRI and video images.. Journal of Phonetics, 2002, 30 (3), pp.533-553. ⟨10.1006/jpho.2002.0166⟩. ⟨hal-00798689⟩
F. Yvon, Philippe Boula de Mareüil, Christophe d’Alessandro, V Aubergé, M Bagein, et al.. Objective evaluation of grapheme to phoneme conversion for text-to-speech synthesis in French. Computer Speech and Language, 1998, 12 (4), pp.393 – 410. ⟨10.1006/csla.1998.0104⟩. ⟨hal-02000962⟩
Communications dans un congrès149 documents
Gérard Bailly, Olivier Perrotin. The GIPSA-Lab Text-To-Speech System for the Blizzard Challenge 2025. Blizzard Challenge 2025 – 19th Workshop, Aug 2025, Gröningen, Netherlands. ⟨10.21437/Blizzard.2025-5⟩. ⟨hal-05368789v2⟩
Gérard Bailly, Elisabeth André, Erica Cooper, Benjamin R. Cowan, Jens Edlund, et al.. Hot topics in speech synthesis evaluation. SSW 2025 – 13th edition of the Speech Synthesis Workshop, Aug 2025, Leeuwarden, Netherlands. pp.1 – 7, ⟨10.21437/ssw.2025-1⟩. ⟨hal-05304059⟩
Sanjana Sankar, Martin Lenglet, Gérard Bailly, Denis Beautemps, Thomas Hueber. Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model. ICASSP 2025 – IEEE International Conference on Acoustics, Speech and Signal Processing, Apr 2025, Hyderabad, India. ⟨hal-04875292⟩
Rami Younes, Frédéric Elisei, Damien Pellier, Gérard Bailly. Impact of verbal instructions and deictic gestures of a cobot on the performance of human coworkers. IEEE-RAS International Conference on Humanoid Robots, Nov 2024, Nancy, France. pp. 1025–1032. ⟨hal-04708887⟩
Fabien Ringeval, Björn Schuller, Gérard Bailly, Safaa Azzakhnini, Hippolyte Fournier. EVAC 2024 – Empathic Virtual Agent Challenge: Appraisal-based Recognition of Affective States. ICMI 2024 – 26th ACM International Conference on Multimodal Interaction, Nov 2024, San Jose, Costa Rica. pp.677 – 683, ⟨10.1145/3678957.3689029⟩. ⟨hal-04984569⟩
Frédéric Elisei, Léa Haefflinger, Gérard Bailly. RoboTrio2: Annotated Interactions of a Teleoperated Robot and Human Dyads for Data-Driven Behavioral Models. HHAI 2024 – Third International Conference on Hybrid Human-Artificial Intelligence (HHAI), Oct 2024, Malmö, Sweden. pp.84-92. ⟨hal-04905371⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. FastLips: an End-to-End Audiovisual Text-to-Speech System with Lip Features Prediction for Virtual Avatars. Interspeech 2024 – 25th Annual Conference of the International Speech Communication Association, Sep 2024, Kos, Greece. pp.3450-3454, ⟨10.21437/Interspeech.2024-462⟩. ⟨hal-04683663⟩
Delphine Charuau, Andrea Briglia, Erika Godde, Gérard Bailly. Training speech-breathing coordination in computer-assisted reading. Interspeech 2024 – 25th Annual Conference of the International Speech Communication Association, Sep 2024, Kos, Greece. pp.5128 – 5132, ⟨10.21437/interspeech.2024-992⟩. ⟨hal-04700395⟩
Erika Godde, Marie-Line Bosse, Gérard Bailly. A reading karaoke to improve reading rate, reading prosody and comprehension. SSSR 2024 – 31st Annual Conference Society for the Scientific Study of Reading, Society for the Scientific Study of Reading, Jul 2024, Copenhague, Denmark. ⟨hal-04664520⟩
Delphine Charuau, Andrea Briglia, Erika Godde, Gérard Bailly. Entraînement de la coordination respiration-parole en apprentissage de la lecture assistée par ordinateur. JEP-TALN-RECITAL 2024 – 35èmes Journées d’Études sur la Parole (JEP 2024) 31ème Conférence sur le Traitement Automatique des Langues Naturelles (TALN 2024) 26ème Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RECITAL 2024), Jul 2024, Toulouse, France. pp.351-360. ⟨hal-04623086⟩
Gérard Bailly, Romain Legrand, Martin Lenglet, Frédéric Elisei, Maëva Garnier, et al.. Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS. LREC-COLING 2024 – Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Calzolari, Nicoletta; Kan, Min-Yen; Hoste, Veronique; Lenci, Alessandro; Sakti, Sakriani; Xue, Nianwen, May 2024, Turin, Italy. pp.5689-5695. ⟨hal-04709149⟩
Léa Haefflinger, Frédéric Elisei, Brice Varini, Gérard Bailly. Probing the Inductive Biases of a Gaze Model for Multi-party Interaction. HRS 2024 – 2024 ACM/IEEE International Conference on Human-Robot Interaction (HRI ’24), Mar 2024, Boulder, CO, United States. pp.507-511, ⟨10.1145/3610978.3640702⟩. ⟨hal-04510252⟩
Léa Haefflinger, Frédéric Elisei, Béatrice Bouchot, Brice Varini, Gérard Bailly. Data-Driven Generation of Eyes and Head Movements of a Social Robot in Multiparty Conversation. ICSR 2023 – 15th International Conference on Social Robotics (ICSR 2023), Dec 2023, Doha, Qatar. pp.191-203, ⟨10.1007/978-981-99-8715-3_17⟩. ⟨hal-04335472⟩
Olivier Perrotin, Brooke Stephenson, Silvain Gerber, Gérard Bailly. The Blizzard Challenge 2023. Blizzard Challenge 2023 – 18th Workshop, Aug 2023, Grenoble, France. pp.1-27, ⟨10.21437/Blizzard.2023-1⟩. ⟨hal-04269927⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. The GIPSA-Lab Text-To-Speech System for the Blizzard Challenge 2023. Blizzard Challenge 2023 – 18th Workshop, Aug 2023, Grenoble, France. pp.34-39, ⟨10.21437/Blizzard.2023-3⟩. ⟨hal-04269935⟩
Gérard Bailly, Martin Lenglet, Olivier Perrotin, Esther Klabbers. Advocating for text input in multi-speaker text-to-speech systems. SSW 2023 – 12th ISCA Speech Synthesis Workshop (SSW2023), Gérard Bailly; Olivier Perrotin; Thomas Hueber; Damien Lolive; Nicolas Obin, Aug 2023, Grenoble, France. pp.1-7, ⟨10.21437/SSW.2023-1⟩. ⟨hal-04257685⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. Local Style Tokens: Fine-Grained Prosodic Representations For TTS Expressive Control. SSW 2023 – 12th ISCA Speech Synthesis Workshop (SSW2023), Gérard Bailly; Olivier Perrotin; Thomas Hueber; Damien Lolive; Nicolas Obin, Aug 2023, Grenoble, France. pp.120-126, ⟨10.21437/SSW.2023-19⟩. ⟨hal-04257713⟩
Léa Haefflinger, Frédéric Elisei, Silvain Gerber, Béatrice Bouchot, Jean-Philippe Vigne, et al.. On the Benefit of Independent Control of Head and Eye Movements of a Social Robot for Multiparty Human-Robot Interaction. HCII 2023 – 25th International Conference on. Human-Computer Interaction HCII 2023, Jul 2023, Copenhague, Denmark. pp.450-466, ⟨10.1007/978-3-031-35596-7_29⟩. ⟨hal-04185780⟩
Hanna Chainay, Eleonore Trâne, Hippolyte Fournier, Olivier Koenig, Franck Tarpin-Bernard, et al.. The value of a virtual assistant to improve engagement in computerized cognitive training at home : An exploratory study. JEV 2023 – Journées d’Etude du Vieillissement, Jun 2023, Tours, France. ⟨hal-04859443⟩
Maria-Loulou Hajj, Martin Lenglet, Olivier Perrotin, Gérard Bailly. Comparing NLP solutions for the disambiguation of French heterophonic homographs for end-to-end TTS systems. SPECOM 2022 – 24th International Conference on Speech and Computer (SPECOM), Nov 2022, Kitt Gurugram, India. pp.265-278, ⟨10.1007/978-3-031-20980-2_23⟩. ⟨hal-03858736⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. Speaking Rate Control of end-to-end TTS Models by Direct Manipulation of the Encoder’s Output Embeddings. Interspeech 2022 – 23rd Annual Conference of the International Speech Communication Association, Sep 2022, Incheon, South Korea. pp.11-15, ⟨10.21437/interspeech.2022-759⟩. ⟨hal-03793220v2⟩
Rami Younes, Gérard Bailly, Damien Pellier, Frédéric Elisei. Automatic Verbal Depiction of a Brick Assembly for a Robot Instructing Humans. SIGDIAL 2022 – 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL 2022), Sep 2022, Edinburgh, United Kingdom. ⟨hal-03754055⟩
Léa Haefflinger, Frédéric Elisei, Gérard Bailly. Multiparty attention management for an embodied conversational agent. JNRH 2022 – Journées Nationales de la Robotique Humanoïde, Jul 2022, Angers, France. ⟨hal-03780683⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. Modélisation de la Parole avec Tacotron2 : Analyse acoustique et phonétique des plongements de caractère. JEP 2022 – 34e Journées d’Études sur la Parole, Jun 2022, Noirmoutier, France. ⟨hal-03727735⟩
Loriane Koelsch, Frédéric Elisei, Ludovic Ferrand, Pierre Chausse, Gérard Bailly, et al.. Impact of social presence of humanoid robots: does competence matter?. ICSR 2021 – International Conference on Remote Sensing, Nov 2021, Singapour, Singapore. ⟨hal-03411321⟩
Loriane Koelsch, Gérard Bailly, Frédéric Elisei, Pascal Huguet, Ludovic Ferrand. L’impact des robots sur notre cognition : l’effet de présence robotique. WACAI 2021 – Workshop sur les “Affects, Compagnons Artificiels et Interactions” (ACAI), Centre National de la Recherche Scientifique [CNRS], Oct 2021, Saint Pierre d’Oléron, France. ⟨hal-03377544⟩
Olivier Perrotin, Hussein El Amouri, Gérard Bailly, Thomas Hueber. Evaluating the Extrapolation Capabilities of Neural Vocoders to Extreme Pitch Values. Interspeech 2021 – 22nd Annual Conference of the International Speech Communication Association, Aug 2021, Brno, Czech Republic. pp.11-15, ⟨10.21437/Interspeech.2021-1547⟩. ⟨hal-03338483⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. Impact of Segmentation and Annotation in French end-to-end Synthesis. SSW 2021 – 11th ISCA Speech Synthesis Workshop, Aug 2021, Budapest, Hungary. pp.13-18, ⟨10.21437/SSW.2021-3⟩. ⟨hal-03362000⟩
Erika Godde, Gérard Bailly, Marie-Line Bosse. Suivi longitudinal de la fluence en lecture par évaluation automatique de la parole. EIAH 2021 – 10e Conférence sur les Environnements Informatiques pour l’Apprentissage Humain, Jun 2021, Fribourg (Virtual), Suisse. pp.70-81. ⟨hal-03292753⟩
Sonia Mandin, Ahmed Zaher, Svetlana Meyer, Mathieu Loiseau, Gérard Bailly, et al.. Expérimentation à grande échelle d’applications pour tablettes pour favoriser l’apprentissage de la lecture et de l’anglais oral. EIAH 2021 – 10e Conférence sur les Environnements Informatiques pour l’Apprentissage Humain, Marie Lefevre, Christine Michel, Jun 2021, Fribourg, Suisse. pp.118-129. ⟨hal-03292798⟩
Sonia Mandin, Sylviane Valdois, Gérard Bailly, Mathieu Loiseau. FLUENCE : projet de conception et d’expérimentation in-situ, longitudinale et à grande échelle d’applications tablettes pour prévenir les difficultés d’apprentissage de la lecture. SILE 2021, May 2021, Sherbrooke, Canada. ⟨hal-03248965⟩
Sonia Mandin, Mathieu Loiseau, Gérard Bailly, Anne Blavot, Marie-Line Bosse, et al.. EVASION, ELARGIR et LUCIOLE : 3 jeux tablettes du projet FLUENCE pour prévenir les difficultés d’apprentissage de la lecture et de l’anglais. PRUNE II 2021 – Colloque Perspectives de Recherches sur les Usages du Numérique dans l’Éducation, Apr 2021, Paris (virtuel), France. ⟨hal-03187547⟩
Sonia Mandin, Mathieu Loiseau, Gérard Bailly, Sylviane Valdois. Évaluation de dispositifs numériques innovants pour l’apprentissage de la lecture et de l’anglais : une expérimentation longitudinale en condition écologique. SFERE 2021 – 2ème édition du Colloque de SFERE-Provence, Mar 2021, Marseille, France. ⟨hal-03187570⟩
Gérard Bailly, Frédéric Elisei. Speech in action: designing challenges that require incremental processing of self and others’ speech and performative gestures. NLG4HRI 2020 – 2nd Workshop on Natural Language Generation for Human-Robot Interaction, Dec 2020, Dublin (virtual), Ireland. ⟨hal-03084920⟩
Gérard Bailly, Erika Godde, Anne-Laure Piat-Marchand, Marie-Line Bosse. Predicting Multidimensional Subjective Ratings of Children’ Readings from the Speech Signals for the Automatic Assessment of Fluency. LREC 2020 – 12th Conference on Language Resources and Evaluation (LREC 2020), May 2020, Marseille, France. pp.317-322. ⟨hal-03039160⟩
Erika Godde, Gérard Bailly, Marie-Line Bosse. Reading Prosody Development: Automatic Assessment for a Longitudinal Study. SLaTE 2019 – 8th ISCA Workshop on Speech and Language Technology in Education, Sep 2019, Graz, Austria. ⟨10.21437/SLaTE.2019-20⟩. ⟨hal-02181469⟩
Erika Godde, Gérard Bailly, Marie-Line Bosse. Un Karaoké pour Entraîner Prosodie et Compréhension en Lecture. EIAH 2019 – Environnements Informatiques pour l’Apprentissage Humain, Jun 2019, Paris, France. ⟨hal-02141164⟩
Omar Mohammed, Gérard Bailly, Damien Pellier. Style Transfer and Extraction for the Handwritten Letters Using Deep Learning. ICAART 2019 – 11th International Conference on Agents and Artificial Intelligence, Feb 2019, Prague, Czech Republic. ⟨hal-02049006⟩
Omar Mohammed, Gérard Bailly, Damien Pellier. Handwriting Styles: Benchmarks and Evaluation Metrics. DTL 2018 – IEEE International Workshop on Deep and Transfer Learning (DTL 2018), Oct 2018, Valencia, Spain. ⟨hal-01900765⟩
Quan Nguyen, Laurent Girin, Gérard Bailly, Frédéric Elisei, Duc-Canh Nguyen. Autonomous Sensorimotor Learning for Sound Source Localization by a Humanoid Robot. IROS 2018 – Workshop on Crossmodal Learning for Intelligent Robotics in conjunction with IEEE/RSJ IROS, Oct 2018, Madrid, Spain. ⟨hal-01921882⟩
Branislav Gerazov, Gérard Bailly, Yi Xu. A Weighted Superposition of Functional Contours Model for Modelling Contextual Prominence of Elementary Prosodic Contours. Interspeech 2018 – 19th Annual Conference of the International Speech Communication Association, Sep 2018, Hyderabad, India. ⟨10.21437/interspeech.2018-1286⟩. ⟨hal-01921906⟩
Duc-Canh Nguyen, Gérard Bailly, Frédéric Elisei. Comparing cascaded LSTM architectures for generating head motion from speech in task-oriented dialogs. HCI 2018 – 20th International Conference on Human-Computer Interaction, Jul 2018, Las Vegas, United States. pp.164-175. ⟨hal-01848063⟩
Gérard Bailly, Frédéric Elisei. Demonstrating and Learning Multimodal Socio-communicative Behaviors for HRI: Building Interactive Models from Immersive Teleoperation Data. FAIM/ISCA Workshop on Artificial Intelligence for Multimodal Human Robot Interaction, Jul 2018, Stockholm, Sweden. pp.39-43, ⟨10.21437/AI-MHRI.2018-10⟩. ⟨hal-01835008⟩
Remi Cambuzat, Frédéric Elisei, Gérard Bailly, Olivier Simonin, Anne Spalanzani. Immersive Teleoperation of the Eye Gaze of Social Robots Assessing Gaze-Contingent Control of Vergence, Yaw and Pitch of Robotic Eyes. ISR 2018 – 50th International Symposium on Robotics, VDE, Jun 2018, Munich, Germany. pp.232-239. ⟨hal-01779633⟩
Branislav Gerazov, Gérard Bailly, Yi Xu. The significance of scope in modelling tones in Chinese. TAL 2018 – Sixth International Symposium on Tonal Aspects of Languages (TAL2018), Jun 2018, Berlin, Germany. pp.183-187, ⟨10.21437/TAL.2018-37⟩. ⟨hal-01834964⟩
Branislav Gerazov, Gérard Bailly. PySFC – A System for Prosody Analysis based on the Superposition of Functional Contours Prosody Model. Speech Prosody 2018 – 9th International Conference on Speech Prosody, Jun 2018, Poznan, Poland. pp.774-778, ⟨10.21437/SpeechProsody.2018-157⟩. ⟨hal-01821214⟩
Erika Godde, Gérard Bailly, David Escudero, Marie-Line Bosse, Estelle Gillet-Perret. Evaluation of Reading Performance of Primary School Children: Objective Measurements vs. Subjective Ratings. WOCCI 2017 – 6th Workshop on Child Computer Interaction, Nov 2017, Glasgow, United Kingdom. ⟨10.21437/WOCCI.2017-4⟩. ⟨hal-01638355⟩
Erika Godde, Gérard Bailly, David Escudero, Marie-Line Bosse, Maryse Bianco, et al.. Improving fluency of young readers: introducing a Karaoke to learn how to breath during a Reading-while-Listening task. SLaTE 2017 – 7th ISCA Workshop on Speech and Language Technology in Education, Aug 2017, Stockholm, Sweden. pp.127-131, ⟨10.21437/SLaTE.2017-22⟩. ⟨hal-01575223⟩
Duc-Canh Nguyen, Gérard Bailly, Frédéric Elisei. An Evaluation Framework to Assess and Correct the Multimodal Behavior of a Humanoid Robot in Human-Robot Interaction. GESPIN 2017 – GEstures and SPeech in INteraction, Aug 2017, Posnan, Poland. ⟨hal-01578713⟩
Duc-Canh Nguyen, Gérard Bailly, Frédéric Elisei. Conducting neuropsychological tests with a humanoid robot: design and evaluation. CogInfoCom 2016 – IEEE International Conference on Cognitive Infocommunications, Oct 2016, Wroclaw, Poland. pp.337-342. ⟨hal-01385666⟩
Adela Barbulescu, Rémi Ronfard, Gérard Bailly. Characterization of Audiovisual Dramatic Attitudes. Interspeech 2016 – 17th Annual Conference of the International Speech Communication Association, Sep 2016, San Francisco, United States. pp.585-589, ⟨10.21437/Interspeech.2016-75⟩. ⟨hal-01337077⟩
Gérard Bailly, Frédéric Elisei, Alexandra Juphard, Olivier Moreaud. Quantitative analysis of backchannels uttered by an interviewer during neuropsychological tests. Interspeech 2016 – 17th Annual Conference of the International Speech Communication Association, Sep 2016, San Francisco, United States. ⟨10.21437/Interspeech.2016-22⟩. ⟨hal-01372819⟩
Maël Pouget, Olha Nahorna, Thomas Hueber, Gérard Bailly. Adaptive Latency for Part-of-Speech Tagging in Incremental Text-to-Speech Synthesis. Interspeech 2016 – 17th Annual Conference of the International Speech Communication Association, Sep 2016, San Francisco, CA, United States. pp.2846 – 2850, ⟨10.21437/Interspeech.2016-165⟩. ⟨hal-01374782⟩
Duc-Canh Nguyen, Frédéric Elisei, Gérard Bailly. Demonstrating to a humanoid robot how to conduct neuropsychological tests. JNRH 2016 – Journées Nationales de la Robotique Humanoïde, LAAS – Toulouse, Jun 2016, Toulouse, France. pp.10-12. ⟨hal-01342349⟩
Omar Mohammed, Gérard Bailly, Damien Pellier. Acquiring Human-Robot Interaction skills with Transfer Learning Techniques. ACM/IEEE International Conference on Human-Robot Interaction, Mar 2016, Vienne, Austria. pp.359 – 360, ⟨10.1145/3029798.3034823⟩. ⟨hal-01490211⟩
François Foerster, Gérard Bailly, Frédéric Elisei. Impact of Iris Size and Eyelids Coupling on the Estimation of the Gaze Direction of a Robotic Talking Head by Human Viewers. Humanoids 2015 – IEEE-RAS 15th International Conference on Humanoid Robots, Nov 2015, Séoul, South Korea. pp.148-153. ⟨hal-01228887⟩
Guillermo Gomez, Carole Plasson, Frédéric Elisei, Frédéric Noël, Gérard Bailly. Qualitative assessment of an immersive teleoperation environment for collaborative professional activities in a « beaming » experiment. EuroVR 2015 – European conference for Virtual Reality and Augmented Reality, Oct 2015, Milan, Italy. 8 p. ⟨hal-01228890⟩
Emilie Gerbier, Gérard Bailly, Marie-Line Bosse. The effect of audio-visual synchronization in reading while listening to texts: An eye-tracking study. ESCOP 2015 – 19th Meetings of the European Society for Cognitive Psychology, Sep 2015, Paphos, Cyprus. ⟨hal-01838760⟩
Adela Barbulescu, Gérard Bailly, Rémi Ronfard, Maël Pouget. Audiovisual Generation of Social Attitudes from Neutral Stimuli. FAAVSP 2015 – 1st Joint Conference on Facial Analysis, Animation and Auditory-Visual Speech Processing, Sep 2015, Vienne, Austria. pp.34-39. ⟨hal-01178056⟩
Maël Pouget, Thomas Hueber, Gérard Bailly, Timo Baumann. HMM Training Strategy for Incremental Speech Synthesis. Interspeech 2015 – 16th Annual Conference of the International Speech Communication Association, ISCA, Sep 2015, Dresden, Germany. pp.1201-1205. ⟨hal-01228889⟩
Emilie Gerbier, Gérard Bailly, Marie-Line Bosse. Using Karaoke to enhance reading while listening: impact on word memorization and eye movements. SLaTE 2015 – ISCA Workshop on Speech and Language Technology in Education, Sep 2015, Leipzig, Germany. pp.59-64. ⟨hal-01192870⟩
Gérard Bailly, Frédéric Elisei, Miquel Sauze. Beaming the Gaze of a Humanoid Robot. Human-Robot Interaction. Workshop on Behavior Coordination between Animals, Humans, and Robots, Mar 2015, Portland, United States. ⟨10.1145/2701973.2701992⟩. ⟨hal-01110288⟩
Gérard Bailly, Alaeddine Mihoub, Christian Wolf, Frédéric Elisei. Learning joint multimodal behaviors for face-to-face interaction: performance & properties of statistical models. Human-Robot Interaction. Workshop on Behavior Coordination between Animals, Humans, and Robots, Mar 2015, Portland, United States. ⟨hal-01110290⟩
Alberto Parmiggiani, Marco Randazzo, Marco Maggiali, Frédéric Elisei, Gérard Bailly, et al.. An articulated talking face for the iCub. Humanoids 2014 – IEEE-RAS International Conference on Humanoid Robots, Nov 2014, Madrid, Spain. ⟨hal-01110293⟩
Adela Barbulescu, Rémi Ronfard, Gérard Bailly, Georges Gagneré, Huseyin Cakmak. Beyond Basic Emotions: Expressive Virtual Actors with Social Attitudes. MIG 2014 – 7th International ACM SIGGRAPH Conference on Motion in Games 2014 (MIG 2014), Nov 2014, Los Angeles, United States. pp.39-47, ⟨10.1145/2668084.2668084⟩. ⟨hal-01064989⟩
Alaeddine Mihoub, Gérard Bailly, Christian Wolf. Modeling Perception-Action Loops: Comparing Sequential Models with Frame-Based Classifiers. HAI 2014 – 2nd International Conference on Human-Agent Interaction, Oct 2014, Tsukuba, Japan. pp.309-314. ⟨hal-01061454⟩
Gérard Bailly, Amélie Martin. Assessing objective characterizations of phonetic convergence. Interspeech 2014 – 15th Annual Conference of the International Speech Communication Association, Sep 2014, Singapour, Singapore. pp.P-19-9. ⟨hal-01067610⟩
Alaeddine Mihoub, Gérard Bailly, Christian Wolf. Modeling sensory-motor behaviors for social robots. WACAI 2014 – Workshop Affect, Compagnon Artificiel, Interaction, Jun 2014, Rouen, France. ⟨hal-01527421⟩
Gérard Bailly, Magalie Ochs, Alexandre Pauchet, Humbert Fiorino. Virtual conversational agents and social robots: converging challenges. WACAI 2014 – Workshop Affect, Compagnon Artificiel, Interaction, Jun 2014, Rouen, France. ⟨hal-01488239⟩
Amélie Rochet-Capellan, Gérard Bailly, Susanne Fuchs. Is breathing sensitive to the communication partner?. Speech Prosody 2014 – 7th International Conference on Speech Prosody, May 2014, Dublin, Ireland. pp.613-618. ⟨hal-01004426⟩
Pierre Badin, Thomas Hueber, Gérard Bailly, Didier Demolin, Françoise Raby. Introduction to the proceedings of the SLaTE 2013 workshop on Speech and Language Technology in Education. SLaTE 2013 – Speech and Language Technology in Education, Aug 2013, Grenoble, France. pp.8-10. ⟨hal-00960360⟩
Adela Barbulescu, Thomas Hueber, Gérard Bailly, Rémi Ronfard. Audio-Visual Speaker Conversion using Prosody Features. AVSP 2013 – 12th International Conference on Auditory-Visual Speech Processing, Aug 2013, Annecy, France. pp.11-16. ⟨hal-00842928⟩
Thomas Hueber, Gérard Bailly, Pierre Badin, Frédéric Elisei. Speaker adaptation of an acoustic-to-articulatory inversion model using cascaded Gaussian mixture regressions. Interspeech 2013 – 14th Annual Conference of the International Speech Communication Association, Aug 2013, Lyon, France. pp.2753-2757. ⟨hal-00851894⟩
Gérard Bailly, Amélie Rochet-Capellan, Coriandre Emmanuel Vilain. Adaptation of respiratory patterns in collaborative reading. Interspeech 2013 – 14th Annual Conference of the International Speech Communication Association, Aug 2013, Lyon, France. pp.1653-1657. ⟨hal-00851890⟩
Thomas Hueber, Atef Ben Youssef, Gérard Bailly, Pierre Badin, Frédéric Elisei. Cross-speaker acoustic-to-articulatory inversion using phone-based trajectory HMM. Interspeech 2012 – 13th Annual Conference of the International Speech Communication Association, Sep 2012, Portland, United States. pp.Tue.SS3.08. ⟨hal-00974347⟩
Gérard Bailly, Cécilia Gouvernayre. Pauses and respiratory markers of the structure of book reading. Interspeech 2012 – 13th Annual Conference of the International Speech Communication Association, Sep 2012, Portland, United States. pp.Thu.O9d.05. ⟨hal-00741667⟩
Thomas Hueber, Gérard Bailly, Bruce Denby. Continuous Articulatory-to-Acoustic Mapping using Phone-based Trajectory HMM for a Silent Speech Interface. Interspeech 2012 – 13th Annual Conference of the International Speech Communication Association, Sep 2012, Portland, United States. pp.Tue.P3c.01. ⟨hal-00741682⟩
Amélie Lelong, Gérard Bailly. Original objective and subjective characterization of phonetic convergence. ISICS 2012: International Symposium on Imitation and Convergence in Speech, Sep 2012, Aix-en-Provence, France. pp.O1:2. ⟨hal-00741686⟩
Thomas Hueber, Atef Ben Youssef, Pierre Badin, Gérard Bailly, Frédéric Elisei. Vizart3D : retour articulatoire visuel pour l’aide à la prononciation. JEP-TALN-RECITAL 2012 – conférence conjointe 29e Journées d’Études sur la Parole, 19e Traitement Automatique des Langues Naturelles, 14e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues, Jun 2012, Grenoble, France. pp.17-18. ⟨hal-00725513⟩
Amélie Lelong, Gérard Bailly. Characterizing phonetic convergence with speaker recognition techniques. LISTA 2012 – The Listening Talker Workshop (LISTA 2012), May 2012, Édimbourg, United Kingdom. pp.28-31. ⟨hal-00695558⟩
Gérard Bailly. Human-Machine Interaction. Mutual attention and accommodation during face-to-face interaction. RESCOM 2012 – Researching Communication at UWS Brain, Behaviour and Computation, Dec 2011, Sydney, Australia. ⟨hal-00652587⟩
Thomas Hueber, Atef Ben Youssef, Pierre Badin, Gérard Bailly, Frédéric Elisei. Articulatory-to-acoustic mapping: application to silent speech interface and visual articulatory feedback. PEVOC 2015 – 11th Pan-European Voice Conference, Aug 2011, Marseille, France. ⟨hal-00640396⟩
Atef Ben Youssef, Thomas Hueber, Pierre Badin, Gérard Bailly. Toward a multi-speaker visual articulatory feedback system. Interspeech 2011 – 12th Annual Conference of the International Speech Communication Association, Aug 2011, Florence, Italy. pp.589-592. ⟨hal-00618781⟩
Gérard Bailly, William-Seamus Barbour. Synchronous reading: learning French orthography by audiovisual training. Interspeech 2011 – 12th Annual Conference of the International Speech Communication Association, Aug 2011, Florence, Italy. pp.1153-1156. ⟨hal-00618780⟩
Thomas Hueber, Pierre Badin, Christophe Savariaux, Coriandre Emmanuel Vilain, Gérard Bailly. Differences in articulatory strategies between silent, whispered and normal speech ? A pilot study using ElectroMagnetic Articulography. ISSP 2011 – 9th International Seminar on Speech Production, Jun 2011, Montreal, Canada. pp.n/c. ⟨hal-00724657⟩
Atef Ben Youssef, Pierre Badin, Gérard Bailly. Improvement of HMM-based acoustic-to-articulatory speech inversion. ISSP 2011 – 9th International Seminar on Speech Production, Jun 2011, Montreal, Canada. pp.n/c. ⟨hal-00724652⟩
Atef Ben Youssef, Thomas Hueber, Pierre Badin, Gérard Bailly, Frédéric Elisei. Toward a speaker-independent visual articulatory feedback system. ISSP 2011 – 9th International Seminar on Speech Production, Jun 2011, Montreal, Canada. pp.n/c. ⟨hal-00724655⟩
Thomas Hueber, Pierre Badin, Gérard Bailly, Atef Ben Youssef, Frédéric Elisei, et al.. Statistical mapping between articulatory and acoustic data. Application to Silent Speech Interface and Visual Articulatory Feedback. P3S – 1st International Workshop on Performative Speech and Singing Synthesis [P3S], Mar 2011, Vancouver, Canada. ⟨hal-00640395⟩
Sascha Fagel, Gérard Bailly, Frédéric Elisei, Amélie Lelong. On the Importance of Eye Gaze in a Face-to-Face Collaborative Task. AFFINE 2010 – 3rd International Workshop on Affective Interaction in Natural Environments, Oct 2010, Florence, Italy. pp.81-85. ⟨hal-00531001⟩
Jean-David Boucher, Jocelyne Ventre-Dominey, Peter Ford Dominey, Sascha Fagel, Gérard Bailly. Facilitative Effects of Communicative Gaze and Speech in Human-Robot Cooperation. AFFINE 2010 – 3rd International Workshop on Affective Interaction in Natural Environments, Oct 2010, Florence, Italy. pp.71-74. ⟨hal-00531002⟩
Atef Ben Youssef, Pierre Badin, Gérard Bailly. Acoustic-to-articulatory inversion in speech based on statistical models. AVSP 2010 – 9th International Conference on Auditory-Visual Speech Processing, Sep 2010, Hakone, Kanagawa, Japan. pp.S8-3. ⟨hal-00508279⟩
Gérard Bailly, Amélie Lelong. Speech dominoes and phonetic convergence. Interspeech 2010 – 11th Annual Conference of the International Speech Communication Association, Sep 2010, Makuhari, Japan. pp.1153-1156. ⟨hal-00523890⟩
Atef Ben Youssef, Pierre Badin, Gérard Bailly. Can tongue be recovered from face? The answer of data-driven statistical models. Interspeech 2010 – 11th Annual Conference of the International Speech Communication Association, Sep 2010, Makuhari, Japan. pp.2002-2005. ⟨hal-00508276⟩
Pierre Badin, Atef Ben Youssef, Gérard Bailly, Frédéric Elisei, Thomas Hueber. Visual articulatory feedback for phonetic correction in second language learning. Interspeech 2010 – 11th Annual Conference of the International Speech Communication Association, Sep 2010, Makuhari, Japan. pp.n.c. ⟨hal-00508272⟩
Sascha Fagel, Gérard Bailly. Speech, gaze and head motion in a face-to-face collaborative task. ESSV 2010 – 21st Conference on Electronic Speech Signal Processing, Sep 2010, Berlin, Germany. ⟨hal-00523906⟩
Panikos Heracleous, Pierre Badin, Gérard Bailly, Norihiro Hagita. Exploiting multimodal data fusion in robust speech recognition. ICME 2010 – IEEE International Conference on Multimedia and Expo, Jul 2010, Singapour, Singapore. in press. ⟨hal-00508288⟩
Atef Ben Youssef, Pierre Badin, Gérard Bailly, Viet-Anh Tran. Méthodes basées sur les HMMs et les GMMs pour l’inversion acoustico-articulatoire en parole. JEP 2010 – 28e Journées d’Etudes sur la Parole, May 2010, Mons, Belgique. pp.249-252. ⟨hal-00508281⟩
Atef Ben Youssef, Viet-Anh Tran, Pierre Badin, Gérard Bailly. HMMs and GMMs based methods in acoustic-to-articulatory speech inversion. RJCP 2009 – 8ème Rencontres des Jeunes Chercheurs en Parole, Nov 2009, Avignon, France. pp.Article 182. ⟨hal-00443662⟩
Atef Ben Youssef, Pierre Badin, Gérard Bailly, Panikos Heracleous. Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. Interspeech 2009 – 10th Annual Conference of the International Speech Communication Association, Sep 2009, Brighton, United Kingdom. pp.2255-2258. ⟨hal-00419227⟩
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Tomoki Toda. Multimodal HMM-based NAM-to-speech conversion. Interspeech 2009 – 10th Annual Conference of the International Speech Communication Association, Sep 2009, Brighton, United Kingdom. pp.656-659. ⟨hal-00419232⟩
Pierre Badin, Frédéric Elisei, Lin Huang, Yuliya Tarabalka, Gérard Bailly. Vision of tongue in augmented speech: contribution to speech comprehension and visual tracking strategies. Speech and Face-to-Face communication – A workshop / Summer School dedicated to the Memory of Christian Benoît, Oct 2008, Grenoble, France. ⟨hal-00337412⟩
Gérard Bailly, Yu Fang, Frédéric Elisei, Denis Beautemps. Retargeting cued speech hand gestures for different talking heads and speakers. AVSP 2008 – 7th International Conference on Auditory-Visual Speech Processing, Sep 2008, Moreton Island, Australia. pp.8. ⟨hal-00342426⟩
Sascha Fagel, Gérard Bailly. From 3-D speaker cloning to text-to-audiovisual speech. AVSP 2008 – 7th International Conference on Auditory-Visual Speech Processing, Sep 2008, Moreton Island, Australia. pp.43-46. ⟨hal-00361888⟩
Gérard Bailly, Antoine Begault, Frédéric Elisei, Pierre Badin. Speaking with smile or disgust: data and models. AVSP 2008 – 7th International Conference on Auditory-Visual Speech Processing, Sep 2008, Moreton Island, Australia. pp.111-116. ⟨hal-00333673⟩
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Christian Jutten. Improvement to a NAM captured whisper-to-speech system. Interspeech 2008 – 9th Annual Conference of the International Speech Communication Association, Sep 2008, Brisbane, Australia. pp.1465-1468. ⟨hal-00333288⟩
Barry-John Theobald, Sascha Fagel, Gérard Bailly, Frédéric Elisei. LIPS2008: Visual speech synthesis challenge. Interspeech 2008 – 9th Annual Conference of the International Speech Communication Association, Sep 2008, Brisbane, Australia. pp.2310-2313. ⟨hal-00333655⟩
Sascha Fagel, Frédéric Elisei, Gérard Bailly. From 3-D speaker cloning to text-to-audiovisual speech. Interspeech 2008 – 9th Annual Conference of the International Speech Communication Association, Sep 2008, Brisbane, Australia. pp.2325. ⟨hal-00361886⟩
Pierre Badin, Yuliya Tarabalka, Frédéric Elisei, Gérard Bailly. Can you « read tongue movements »?. Interspeech 2008 – 9th Annual Conference of the International Speech Communication Association, Sep 2008, Brisbane, Australia. pp.2635-2637. ⟨hal-00333688⟩
Gérard Bailly, Oxana Govokhina, Gaspard Breton, Frédéric Elisei, Christophe Savariaux. The trainable trajectory formation model TD-HMM parameterized for the LIPS 2008 challenge. Interspeech 2008 – 9th Annual Conference of the International Speech Communication Association, Sep 2008, Brisbane, Australia. pp.561. ⟨hal-00339043⟩
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Christian Jutten. Amélioration de la conversion de voix chuchotée enregistrée par capteur NAM vers la voix audible. JEP 2008 – 27e Journées d’Etudes sur la Parole, Jun 2008, Avignon, France. pp.110-113. ⟨hal-00339058⟩
Gérard Bailly, Alex Bartroli. Generating Spanish intonation with a trainable prosodic model. Speech Prosody 2008 – 4th International Conference on Speech Prosody, May 2008, Campinas, Brazil. pp.63-66. ⟨hal-00339046⟩
Viet-Anh Tran, Gérard Bailly, Hélène Loevenbruck, Tomoki Toda. Predicting F0 and voicing from NAM-captured whispered speech. Speech Prosody 2008 – 4th International Conference on Speech Prosody, May 2008, Campinas, Brazil. pp.107-110. ⟨hal-00333290⟩
Maxime Berar, Michel Desvignes, Gérard Bailly, Yohan Payan. Reconstruction faciale 3D à partir d’images 3D. RFIA 2008 – 16ème congrès francophone AFRIF-AFIA Reconnaissance des formes et Intelligence Artificielle, Jan 2008, Amiens, France. pp.article 94. ⟨hal-00419243⟩
Yuliya Tarabalka, Pierre Badin, Gérard Bailly, Frédéric Elisei. Peut-on lire sur la langue ? Évaluation de l’apport de la vision de la langue à la compréhension de la parole. JPC 2007 – 2èmes Journées de Phonétique Clinique, Dec 2007, Grenoble, France. ⟨hal-00175679⟩
Yuliya Tarabalka, Pierre Badin, Frédéric Elisei, Gérard Bailly. Can you « read tongue movements »? Evaluation of the contribution of tongue display to speech understanding. ASSISTH 2007 – 1ère Conférence internationale sur l’accessibilité et les systèmes de suppléance aux personnes en situation de handicaps, Nov 2007, Toulouse, France. pp.187-193. ⟨hal-00175680⟩
Antoine Picot, Gérard Bailly, Frédéric Elisei, Stephan Raidt. Scrutinizing natural scenes: controlling the gaze of an embodied conversational agent. IVA 2007 – 7th International Conference on Intelligent Virtual Agents, Sep 2007, Paris, France. pp.50-61. ⟨hal-00170337⟩
Stephan Raidt, Gérard Bailly, Frédéric Elisei. Analyzing and modeling gaze during face-to-face interaction. IVA 2007 – 7th International Conference on Intelligent Virtual Agents, Sep 2007, Paris, France. pp.100-101. ⟨hal-00169579⟩
Stephan Raidt, Gérard Bailly, Frédéric Elisei. Mutual gaze during face-to-face interaction. AVSP 2007 – 6th International Conference on Auditory-Visual Speech Processing, Aug 2007, Hilvarenbeek, Netherlands. pp.180-185. ⟨hal-00169566⟩
Frédéric Elisei, Gérard Bailly, Alix Casari, Stephan Raidt. Towards eyegaze-aware analysis and synthesis of audiovisual speech. AVSP 2007 – 6th International Conference on Auditory-Visual Speech Processing, Aug 2007, Hilvarenbeek, Netherlands. pp.50-56. ⟨hal-00169556⟩
Sascha Fagel, Gérard Bailly, Frédéric Elisei. Intelligibility of natural and 3D-cloned German speech. AVSP 2007 – 6th International Conference on Auditory-Visual Speech Processing, Aug 2007, Hilvarenbeek, Netherlands. pp.56-61. ⟨hal-00169563⟩
Oxana Govokhina, Gérard Bailly, Gaspard Breton. Learning optimal audiovisual phasing for a HMM-based control model for facial animation. SSW2007 – 6th ISCA Workshop on Speech Synthesis (SSW6), Aug 2007, Bonn, Germany. pp.1-4. ⟨hal-00169576⟩
Gérard Bailly, Frédéric Elisei, Stephan Raidt, Alix Casari, Antoine Picot. Embodied conversational agents : computing and rendering realistic gaze patterns. Pacific Rim Conference on Multimedia Processing, Oct 2006, Hangzhou, China. pp.9-18. ⟨hal-00143624⟩
Stephan Raidt, Gérard Bailly, Frédéric Elisei. Plateforme expérimentale de capturerestitution croisée pour l’étude de la communication face-à-face. WACA 2006 – Workshop sur les Agents Conversationnels Animés, Oct 2006, Toulouse, France. pp.C18. ⟨hal-00480347⟩
Antoine Picot, Gérard Bailly, Frédéric Elisei, Stephan Raidt. Scrutation de scènes naturelles par un agent conversationnel animé. WACA 2006 – Workshop sur les Agents Conversationnels Animés, Oct 2006, Toulouse, France. pp.C16. ⟨hal-00480354⟩
Alix Casari, Frédéric Elisei, Gérard Bailly, Stephan Raidt. Contrôle du regard et des mouvements des paupières d’une tête parlante virtuelle. WACA 2006 – Workshop sur les Agents Conversationnels Animés, Oct 2006, Toulouse, France. pp.C15. ⟨hal-00480356⟩
Guillaume Gibert, Gérard Bailly, Frédéric Elisei. Evaluating a virtual speech cuer. Interspeech 2006 – ICSLP, Ninth International Conference on Spoken Language Processing, Sep 2006, Pittsburgh, United States. pp.2430-2433. ⟨hal-00366491⟩
Oxana Govokhina, Gérard Bailly, Gaspard Breton, Paul Bagshaw. TDA: A new trainable trajectory formation system for facial animation. Interspeech 2006 – ICSLP, Ninth International Conference on Spoken Language Processing, Sep 2006, Pittsburgh, United States. pp.2474-2477. ⟨hal-00366489⟩
Gérard Bailly, Ian Gorisch. Generating German intonation with a trainable prosodic model. Interspeech 2006 – ICSLP, Ninth International Conference on Spoken Language Processing, Sep 2006, Pittsburgh, United States. pp.2366-2369. ⟨hal-00366490⟩
Aurélie Clodic, Sara Fleury, Rachid Alami, Raja Chatila, Gérard Bailly, et al.. Rackham: An Interactive Robot-Guide. ROMAN 2006 – IEEE International Workshop on Robots and Human Interactive Communications, Sep 2006, Hatfield, United Kingdom. pp.502-509. ⟨hal-00480381⟩
Oxana Govokhina, Gérard Bailly, Gaspard Breton, Paul Bagshaw. A new trainable trajectory formation system for facial animation. Workshop on Experimental Linguistics, Aug 2006, Athens, Greece. pp.25-32. ⟨hal-00366487⟩
Guillaume Gibert, Gérard Bailly, Frédéric Elisei. Evaluation of a virtual speech cuer. Workshop on Experimental Linguistics, Aug 2006, Athens, Greece. pp.141-144. ⟨hal-00366488⟩
Oxana Govokhina, Gérard Bailly, Gaspard Breton, Paul Bagshaw. Evaluation de systèmes de génération de mouvements faciaux. JEP 2006 – XXVIèmes Journées d’´Etudes sur la Parole, Jun 2006, Rennes, France. pp.305-308. ⟨hal-00366538⟩
Gérard Bailly, Cléo Baras, Patrick Bas, Séverine Baudry, Denis Beautemps, et al.. ARTUS : calcul et tatouage audiovisuel des mouvements d’un personnage animé virtuel pour l’accessibilité d’émissions télévisuelles aux téléspectateurs sourds comprenant la Langue Française Parlée Complétée. HANDICAP 2006, Jun 2006, Paris, France. pp.265-270. ⟨hal-00366492⟩
Guillaume Gibert, Gérard Bailly, Frédéric Elisei. Evaluation d’un système de synthèse 3D de Langue française Parlée Complétée. JEP 2006 – XXVIèmes Journées d’´Etudes sur la Parole, Jun 2006, Rennes, France. pp.495-498. ⟨hal-00366539⟩
Marie-Neige Garcia, Christophe d’Alessandro, Gérard Bailly, Philippe Boula de Mareüil, Michel Morel. A joint prosody evaluation of French text-to-speech synthesis systems. LREC 2006 – 5th edition of the International Conference on Language Ressources and Evaluation, May 2006, Gênes, Italy. pp.307-310, ⟨10.63317/54gksgjyssos⟩. ⟨hal-00103557⟩
Philippe Boula de Mareüil, Christophe d’Alessandro, Alexander Raake, Gérard Bailly, Marie-Neige Garcia, et al.. A joint intelligibility evaluation of French text-to-speech synthesis systems: the EvaSy SUS/ACR campaign. LREC 2006 – 5th edition of the International Conference on Language Ressources and Evaluation, May 2006, Gênes, Italy. pp.2034-2037, ⟨10.63317/2ym3gony58av⟩. ⟨hal-00103571⟩
Stephan Raidt, Gérard Bailly, Frédéric Elisei. Does a Virtual Talking Face Generate Proper Multimodal Cues to Draw User’s Attention to Points of Interest?. LREC 2006 – 5th edition of the International Conference on Language Ressources and Evaluation, May 2006, Gênes, Italy. pp.2544-2549. ⟨hal-00366537⟩
P. Gacon, Pierre-Yves Coulon, Gérard Bailly. Audiovisual Speech Enhancement Experiments for Mouth Segmentation Evaluation. 2006. ⟨hal-00114114⟩
Gérard Bailly, Frédéric Elisei, Pierre Badin, Christophe Savariaux. Degrees of freedom of facial movements in face-to-face conversational speech. International Workshop on Multimodal Corpora, 2006, Genoa, Italy. pp.33-36. ⟨hal-00195551⟩
Stephan Raidt, Gérard Bailly, Frédéric Elisei. Basic components of a face-to-face interaction with a conversational agent: multual attention and deixis. Smart Objects and Ambient Intelligence, Oct 2005, Grenoble, France. pp.247-252. ⟨hal-00366540⟩
P. Gacon, Pierre-Yves Coulon, Gérard Bailly. Modèle statistique et description locale d’apparence non linéaire pour la détection des contours des lèvres. GRETSI 2005 – 20ème colloque GRETSI sur le Traitement du Signal et des Images, GRETSI 2005, Sep 2005, Louvain-la-Neuve, Belgique. ⟨hal-00378354⟩
Philippe Boula de Mareüil, Christophe d’Alessandro, Gérard Bailly, Frédéric Bechet, Marie-Neige Garcia, et al.. Evaluating the pronunciation of proper names by four French grapheme-to-phoneme converters. Interspeech’2005 – Eurospeech. 9th European Conference on Speech Communication and Technology, Sep 2005, Lisbonne, Portugal. pp.1521-1524, ⟨10.21437/Interspeech.2005-447⟩. ⟨hal-00103607⟩
P. Gacon, Pierre-Yves Coulon, Gérard Bailly. Non-Linear Active Model for Mouth Inner and Outer Contours Detection. EUSIPCO 2005 – 13th European Signal Processing Conference (EUSIPCO’05), Sep 2005, Antalya, Turkey. ⟨hal-00378352⟩
Gérard Bailly, Frédéric Elisei, Stephan Raidt. Multimodal face-to-face interaction with a talking face: mutual attention and deixis. Human-Computer Interaction, Jul 2005, France. 10 p. ⟨hal-00516324⟩
Maxime Berar, Michel Desvignes, Gérard Bailly, Yohan Payan. 3D statistical facial reconstruction. 2005, pp.365-370. ⟨hal-00108498⟩
P. Gacon, Pierre-Yves Coulon, Gérard Bailly. Statistical Active Model for Mouth Components Segmentation. ICASSP, 2005, Philadelphia, United States. ⟨hal-00378353⟩
Frédéric Elisei, Gérard Bailly, Guillaume Gibert, Rémi Brun. Capturing data and realistic 3D models for cued speech analysis and audiovisual synthesis. Auditory-Visual Speech Processing Workshop, 2005, Vancouver, Canada. pp.125-130. ⟨hal-00419311⟩
Stephan Raidt, Frédéric Elisei, Gérard Bailly. Face-to-face interaction with a conversationnal agent: eye-gaze and deixis. International Conference on Autonomous Agents and Multiagent Systems, 2005, Utrecht, Netherlands. pp.17-22. ⟨hal-00419299⟩
Lionel Reveret, Gérard Bailly, Pierre Badin. MOTHER: A new generation of talking heads providing a flexible articulatory control for video-realistic speech animation. Int. Conference of Spoken Language Processing, ICSLP’2000, Oct 2000, Pekin, China. ⟨inria-00389362⟩
Poster de conférence8 documents
Cynthia Boggio, Alexis Favre-Felix, Erika Godde, Gérard Bailly, Andrea Briglia, et al.. Le bouquet Fluence : 4 applications numériques pédagogiques pour l’apprentissage des fondamentaux, la lecture et l’anglais. Journée R&T Cognition, Jan 2025, Paris, France. 2025. ⟨hal-04988088⟩
Andrea Briglia, Erika Godde, Delphine Charuau, Marie-Line Bosse, Gérard Bailly. A karaoke-based game to improve the ability of French primary school pupils in planning pauses and breathing while reading aloud. SSSR 2024 – 31st Annual Conference Society for the Scientific Study of Reading, Jul 2024, Copenhague, Denmark. . ⟨hal-04664492⟩
Hippolyte Fournier, Sina Alisamir, Safaa Azzakhnini, Hanna Chainay, Olivier Koenig, et al.. Appraisal-based affect recognition in healthcare: Insights from the THERADIA WoZ corpus. Collectif Cognitif 2024 – 1er Colloque du Collectif Cognitif, May 2024, Paris, France. . ⟨hal-04752973⟩
Martin Lenglet, Olivier Perrotin, Gérard Bailly. A Closer Look at Latent Representations of End-to-end TTS Models. Journée commune AFIA-TLH / AFCP – “Extraction de connaissances interprétables pour l’étude de la communication parlée”, Dec 2023, Avignon, France. . ⟨hal-04269953⟩
Andrea Briglia, Erika Godde, Cynthia Boggio, Marie-Line Bosse, Gérard Bailly. ELARGIR : s’entraîner à lire avec fluence et expressivité. RJCP 2023 – 10ème Rencontres des Jeunes Chercheurs en Parole, Nov 2023, Grenoble, France. . ⟨hal-04328841⟩
Erika Godde, Gérard Bailly, Marie-Line Bosse. Un Karaoké pour Entrainer la Prosodie en Lecture. EIAH 2019 – Environnements Informatiques pour l’Apprentissage Humain, Jun 2019, Paris, France. ⟨hal-02141234⟩
Remi Cambuzat, Frédéric Elisei, Gérard Bailly. SGCS: Stereo Gaze Contingent Steering for Immersive Telepresence. JJCR 2017 – Journée des Jeunes Chercheurs en Robotique, Nov 2017, Bidart, France. , pp.1. ⟨hal-01677481⟩
Remi Cambuzat, Frédéric Elisei, Gérard Bailly. SGCS : Stereo Gaze Contingent Steering for Immersive Telepresence. ECEM 2017 – 19th European Conference on Eye Movements, Aug 2017, Wuppertal, Germany. , European Conference on Eye Movements (ECEM). ⟨hal-01638383⟩
Ouvrages (y compris édition critique et traduction)4 documents
Gérard Bailly, Sylvie Pesty (Dir.). Cognition, Affects et Interaction. 2017, Cognition, Affects et Interaction. ⟨hal-01483705⟩
Gérard Bailly, Pascal Perrier, Eric Vatikiotis-Bateson. Audiovisual Speech Processing. G. Bailly, P. Perrier and E. Vatikiotis-Bateson (Eds.). Cambridge University Press, pp.506, 2012, 978-1107006829. ⟨hal-00691938⟩
Marion Dohen, Jean-Luc Schwartz, Gérard Bailly (Dir.). Speech and face-to-face communication. Elsevier, pp.135, 2010, ISSN:0167-6393. ⟨hal-00941200⟩
Chapitres d’ouvrage16 documents
Loriane Koelsch, Frédéric Elisei, Ludovic Ferrand, Pierre Chausse, Gérard Bailly, et al.. Impact of social presence of humanoid robots: does competence matter?. Social Robotics. ICSR 2021, 13086, Springer International Publishing, pp.729-739, 2021, Lecture Notes in Computer Science, 978-3-030-90524-8. ⟨10.1007/978-3-030-90525-5_64⟩. ⟨hal-05541855⟩
Franck Tarpin-Bernard, Joan Fruitet, Jean-Philippe Vigne, Patrick Constant, Hanna Chainay, et al.. THERADIA: Digital Therapies Augmented by Artificial Intelligence. Advances in Neuroergonomics and Cognitive Engineering. AHFE 2021., 259, Springer International Publishing, pp.478-485, 2021, Lecture Notes in Networks and Systems, ISBN 978-3-030-80284-4. ⟨10.1007/978-3-030-80285-1_55⟩. ⟨hal-03411225⟩
Gérard Bailly, Alaeddine Mihoub, Christian Wolf, Frédéric Elisei. Gaze and face-to-face interaction. Geert Brône & Bert Oben. Eye-tracking in Interaction. Studies on the role of eye gaze indialogue, Benjamins, pp.139 – 168, 2018, ⟨10.1075/ais.10.07bai⟩. ⟨hal-01939223⟩
Alaeddine Mihoub, Gérard Bailly, Christian Wolf. Social behavior modeling based on Incremental Discrete Hidden Markov Models. Human Behavior Understanding. 4th International Workshop, HBU 2013, Barcelona, Spain, October 22, 2013. Proceedings, Springer International Publishing, pp.172-183, 2013, Lecture Notes in Computer Science, n°8212, 978-3-319-02714-2. ⟨10.1007/978-3-319-02714-2_15⟩. ⟨hal-00851903⟩
Gérard Bailly, Pierre Badin, Lionel Reveret, Atef Ben Youssef. Sensorimotor characteristics of speech production. G. Bailly, P. Perrier and E. Vatikiotis-Bateson (Eds.). Audiovisual Speech Processing, Cambridge University Press, pp.368-396, 2012, 978-1107006829. ⟨hal-00694313⟩
Amélie Lelong, Gérard Bailly. Study of the phenomenon of phonetic convergence thanks to speech dominoes. A. Esposito, A. Vinciarelli, K. Vicsi, C. Pelachaud and A. Nijholt. Analysis of Verbal and Nonverbal Communication and Enactment: The Processing Issue, Springer Verlag, pp.280-293, 2011, LNCS AI. ⟨hal-00603164⟩
Gérard Bailly, Frédéric Elisei, Stephan Raidt. Des machines parlantes aux agents conversationnels incarnés. Catherine Garbay, Daniel Kayser (Eds.). Informatique et Sciences Cognitives : influences ou confluences?, Ophrys, pp.215-234, 2011. ⟨hal-00634982⟩
Sascha Fagel, Gérard Bailly. Speech, Gaze and Head Motion in a Face-to-Face Collaborative Task. Anna Esposito, Antonietta M. Esposito, Raffaele Martone, Vincent C. Müller, Gaetano Scarpetta. Toward Autonomous, Adaptive, and Context-Aware Multimodal Interfaces: Theoretical and Practical Issues, Springer-Verlag, pp.265-274, 2010, Lecture Notes in Computer Science (LNCS) n°6456, 978-3-642-18183-2. ⟨hal-00531003⟩
Gérard Bailly, Pierre Badin, Denis Beautemps, Frédéric Elisei. Speech technologies for augmented communication. J. Mullennix and S. Stern. Computer synthesized speech technologies: tools for aiding impairment, IGI Global, Hershey, PA, pp.116-128, 2010. ⟨hal-00473026⟩
Virginie Attina, Guillaume Gibert, Marie-Agnès Cathiard, Gérard Bailly, Denis Beautemps. The Analysis of French Cued Speech Production-Perception. Towards a complete Text-to-Cued Speech Synthesizer. C. LaSasso, K. L. Crain and J. Leybaert. Cued Speech and Cued Language Development for Deaf and Hard of Hearing Children, Plural Publishing Inc., San Diego, CA,, pp.449-466, 2010, 978-1-59756-334-5. ⟨hal-00473028⟩
Gérard Bailly, Catherine Pelachaud. Parole et expression des émotions sur le visage d’humanoïdes virtuels. P. Fuchs, G. Moreau and S. Donikian. Traité de la réalité virtuelle: Volume 5 : les humains virtuels, Presses de l’Ecole des Mines de Paris, pp.187-208, 2009, Mathématique et informatique, 9782911256042. ⟨hal-00376537⟩
Pierre Badin, Frédéric Elisei, Gérard Bailly, Yuliya Tarabalka. An audiovisual talking head for augmented speech generation: models and animations based on a real speaker’s articulatory data. F.J. Perales & R.B. Fisher. Proceedings of the Vth Conference on Articulated Motion and Deformable Objects (AMDO 2008), 5098, Springer Verlag: Berlin, Heidelberg, Germany, pp.132-143, 2008, Lecture Notes in Computer Science, 5098, ⟨10.1007/978-3-540-70517-8_14⟩. ⟨hal-00296599⟩
Christophe d’Alessandro, Philippe Boula de Mareüil, Marie-Neige Garcia, Gérard Bailly, Michel Morel, et al.. La campagne EvaSy d’évaluation de la synthèse de la parole à partir du texte. Stéphane Chaudiron; Khalid Choukri. L’ évaluation des technologies de traitement de la langue : les campagnes Technolangue, Hermès science publ.; Lavoisier, pp.183-208, 2008, (IC2. Cognition et traitement de l’information), 978-2-7462-1992-2. ⟨hal-00361911⟩
Gérard Bailly, Frédéric Elisei, Stephan Raidt. Virtual talking heads and ambiant face-to-face communication. A. Esposito; E. Keller; M. Marinaro and M. Bratanic. The fundamentals of verbal and non-verbal communication and the biometrical issue, IOS Press BV, Amsterdam, pp.302-316, 2007, 978-1-58603-733-8. ⟨hal-00361913⟩
Maxime Berar, Gérard Bailly, Matthieu Chabanas, Michel Desvignes, Frédéric Elisei, et al.. Morphing Generic Organs To Speaker-Specific Anatomies. J. Harrington & M. Tabain. Speech Production: Models, Phonetic Processes, and Techniques, Psychology Press, New York, pp.341-362, 2006, Chapter 20, ISBN 1841694371. ⟨hal-00108522⟩
Gérard Bailly. Caracterisation of formant trajectories by tracking vocal tract resonances. Christel Sorin; Joseph Mariani; Henry Meloni; Jean Schoentgen. Levels in Speech Communication: Relations and Interactions, Elsevier, pp.91-102, 1995, 978-0444818461. ⟨hal-04798673⟩
Brevets1 document
Oxana Govokhina, Gaspard Breton, Gérard Bailly. Procédé de repositionnement des frontières de phonèmes pour la synthèse visuelle des mouvements faciaux liés à la parole. France, N° de brevet: FR0757063. Département Parole et Cognition. 2008. ⟨hal-00361889⟩