Published on Sat Oct 24 2020

Open-Domain Dialogue Generation Based on Pre-trained Language Models

Yan Zeng, Jian-Yun Nie

Pre-trained language models have been successfully used in response generation for open-domain dialogue. Four main frameworks have been proposed: Transformer-ED using Transformer encoder and decoder separately for source and target sentences.

0
0
0
Abstract

Pre-trained language models have been successfully used in response generation for open-domain dialogue. Four main frameworks have been proposed: (1) Transformer-ED using Transformer encoder and decoder separately for source and target sentences; (2) Transformer-Dec using Transformer decoder for both source and target sentences; (3) Transformer-MLM using Transformer decoder that applies bi-directional attention on the source side and left-to-right attention on the target side with masked language model objective; and (4) Transformer-AR that uses auto-regressive objective instead. In this study, we compare these frameworks on 3 datasets, and our comparison reveals that the best framework uses bidirectional attention on the source side and does not separate encoder and decoder. We also examine model discrepancy, and our experiments confirm that the performance of a model is directly impacted by the underlying discrepancies. We then propose two correction methods to reduce the discrepancies, and both improve the model performance. These results show that discrepancies is an important factor to consider when we use a pre-trained model, and a reduction in discrepancies can lead to improved performance.

Thu Oct 15 2020
NLP
Pretrained Language Models for Dialogue Generation with Multiple Input Sources
Large-scale pretrained language models have achieved outstanding performance on natural language understanding tasks. We explore various methods to fuse multiple separate attention information corresponding to different sources. Our experimental results show that proper fusion methods deliver higher relevance with dialogue history than simple fusion baselines.
0
0
0
Mon Mar 09 2020
Artificial Intelligence
An Empirical Investigation of Pre-Trained Transformer Language Models for Open-Domain Dialogue Generation
We present an empirical investigation of pre-trained Transformer-based auto-regressive language models for the task of open-domain dialogue generation. Corpora of News and Wikipedia in Chinese and English are collected for the pre-training stage respectively.
0
0
0
Thu Oct 17 2019
NLP
PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variable
Pre-training models have been proved effective for a wide range of natural language processing tasks. We propose a novel dialogue generation pre-training framework to support various kinds of conversations.
0
0
0
Tue Feb 09 2021
Artificial Intelligence
AuGPT: Dialogue with Pre-trained Language Models and Data Augmentation
Attention-based pre-trained language models such as GPT-2 brought.considerable progress to end-to-end dialogue modelling. However, they also.present considerable risks for task-oriented dialogue, such as lack of.knowledge grounding or diversity. To address these issues,
0
0
0
Sun Sep 05 2021
NLP
SideControl: Controlled Open-domain Dialogue Generation via Additive Side Networks
The SideControl framework leverages a novel control attributes loss to incorporate useful control signals. It is shown to perform well with very limited training samples. Results show that it has better controllability, higher generation quality and better sample efficiency.
0
0
0
Thu Apr 30 2020
NLP
Towards Unsupervised Language Understanding and Generation by Joint Dual Learning
In modular dialogue systems, natural language understanding (NLU) and natural language generation (NLG) are two critical components. However, the property between understanding and generation has been rarely explored. The proposed approach is capable of boosting the performance of both NLU and NLG.
0
0
0