Job Search

Recruit Detail

Find out more about the work we do, the experience and skills we can bring to the table, and our terms and conditions.

Company Name

Honda Research Institute Japan Co., Ltd.

Job Type

Research Position - Research and Development of Multimodal Dialogue Technology in Human-Centered AI

Work Detail

Established in 2003 in Japan, the United States, and Europe, as a wholly owned subsidiary of Honda R&D Co., Ltd., HRI aims to be an organization that challenges new fields beyond automotive technology. Our guiding principle is "Innovate through Science." Towards the realization of a hybrid society where people, the global environment, and intelligent systems coexist in harmony, we focus on communication and sensing, such as Cooperative Intelligence and Cooperative Devices, through a multifaceted approach drawing from artificial intelligence, robotics, systems science, neuroscience, materials science, psychology, and social ethics. We are engaged in various research projects. * Research to realize multimodal communication between humans and machines through interaction that considers human mental states. * Development of a speech recognition system that promotes mutual understanding between people with hearing impairments and those with normal hearing. * Development of a system for estimating a person's internal state (e.g., level of understanding) using cameras, wearable sensors, etc. ■ Job Openings We are seeking advanced researchers/engineers specializing in multimodal dialogue technology in human-centered AI to drive the advancement of next-generation dialogue systems. [Job Details] * Advanced Multimodal Architecture Focus on strengthening internal inference (Chain-of-Thought) and reducing latency to design and verify novel research concepts for next-generation dialogue models. * Development of Contextual Understanding Models Analyze continuous audio, video, and text data to build algorithms that adapt to user behavior and environmental context. * Solving Real-World Problems Design a system to convert real-world interaction signals into high-quality training data, enabling continuous model improvement and bias reduction. * Promoting Global Collaboration Collaborate with overseas sister research institutes (Germany and the USA) to tackle research themes that lead to business opportunities. * Technical Leadership and Growth Responsible for creating new research themes, project management, nurturing young researchers, and presenting at international conferences. Expected to take on a team leader role in the future. [Job Characteristics] You will proactively handle the entire research lifecycle, from research conception to implementation, evaluation, and practical application. As a role connecting academic research and real-world applications, you will gain experience in elevating foundational models into systems that can be practically operated. With a high degree of autonomy, you can propose and promote your own research themes and actively present at top-level international conferences. • Multimodal Fusion/LLM Foundation models such as Vision-Language Model (VLM), cross-modal attention • Human-centered computer vision Speaker detection, gaze estimation, posture estimation, facial expression/attribute analysis • Speech/natural language processing Speaker recognition/speech recognition integrated with visual information, dialogue context understanding Based on Honda's philosophy of "Technology for people," we aim to create unique technologies for the world with speed and originality. Based on the latest trends in multimodal AI, we identify high-value research themes and create results through rapid prototyping, experimentation, and verification. Research results will be actively disseminated through international conferences and papers. You will promote work in collaboration with team members and several researchers from external research institutions. One of the attractions is the opportunity to work globally in collaboration with our overseas sister companies (Germany, USA). Research projects are driven under the Research Division Manager (Department: Research Division). Research projects are proposed by the individual, including content and budget, and are launched upon board approval. • A highly international workplace with approximately half of all employees being foreign nationals. Communication within the company is conducted daily in both Japanese and English. • The workplace is located within Honda Motor Co., Ltd.'s "Wako Campus," which has received numerous awards as a suburban office. • Researchers have a great deal of autonomy, and the external publication of research results is actively encouraged. There is a culture that respects the company's direction while also valuing the will of the researchers.

Ideal Profile

[MUST] * Master's or PhD in Machine Learning/AI/Computer Science (equivalent practical experience acceptable) * Knowledge of deep learning and multimodal processing spanning NLP, CV, and speech processing * Experience implementing Python and PyTorch/TensorFlow, etc. * 3+ years of practical/research experience in related fields * High level of teamwork and business-level English proficiency [WANT] * Experience with multimodal learning, LLM, and conversational AI * Publication record at top conferences (AAAI, NeurIPS, ICASSP, etc.) * Experience using OSS tools (ESPnet, Hugging Face, etc.) * Experience training and operating large-scale models * Collaborative research experience with overseas research institutions

Work Location

Wako City, Saitama Prefecture

Phd. Stating Salary

Expected annual salary: 6.5 million to 12 million yen

Selection Flow

After a document review, candidates receive a job offer following two interviews. Additionally, candidates are required to take the SPI test and undergo a reference check before receiving a job offer.

Similar Recruits

CyberAgent, Inc.

Job Type
[AI Lab] (Research Engineer) We are looking for a research engineer in the field of voice dialogue!

Working in collaboration with research scientists and business units, you will leverage the knowledge accumulated through field testing and R&D in real-world environments such as commercial facilities and accommodations to implement an autonomous voice dialogue system that understands human behavior and speech and engages in appropriate voice interaction. You will then verify its effectiveness in real-world settings. In the future, as a technical specialist, you will develop a foundational system that can be used across a wide range of fields, while understanding the organization's challenges. ▼Main Tasks Development of a voice dialogue system using a dialogue architecture currently under research and development at the AI ​​Lab Implementation of recognition functions (image processing and voice processing) based on robot-human interaction Implementation of machine learning models on edge devices Construction of a system for deploying robot recognition function components to the edge via the cloud Discussion with researchers, proposal, implementation, and introduction of technologies and models Construction of a system for collecting analytical data for research Collaboration with industry-academia partners and joint research institutions, incorporating academic knowledge, and applying it to product development, problem-solving, and research paper analysis.