Skip to main content

Study notes: Co-Attentive Multi-Task Learning for Explainable Recommendation


This paper is a direct extension of this paper. It adds an explanation task objective and jointly training both rating prediction and explanation tasks.

Please refer my other post for multi-pointer co-attention learning.

Figure 1: Co-Attentive Multi-Task Learning for Explainable Recommendation
From Figure 1, one major difference from this post is Task 2. Task 2 itself is a GRU network used to generate text explanations. Let's denote its output $\boldsymbol{o}_t$ as the distribution of corresponding words. $Y=(y_1, ..., y_T)$ the generated texts.

There are two additional losses from Task 2.

1. Concept relevance loss $\mathcal{L}_c$. During training, $\mathcal{L}_c$ is used to increase the probability that the selected concepts appear in $Y$. It is computed by

2. Negative log-likelihood loss $\mathcal{L}_n$. To ensure that the generated words are similar to the ground truth ones.
Plus, the original rating prediction loss:

The model jointly training the following objective:

Comments

Popular posts from this blog

Reading Notes: Probabilistic Model-Agnostic Meta-Learning

Probabilistic Model-Agnostic Meta-Learning Reading Notes: Probabilistic Model-Agnostic Meta-Learning This post is a reading note for the paper "Probabilistic Model-Agnostic Meta-Learning" by Finn et al. It is a successive work to the famous MAML paper , and can be viewed as the Bayesian version of the MAML model. Introduction When dealing with different tasks of the same family, for example, the image classification family, the neural language processing family, etc.. It is usually preferred to be able to acquire solutions to complex tasks from only a few samples given the past knowledge of other tasks as a prior (few shot learning). The idea of learning-to-learn, i.e., meta-learning, is such a framework. What is meta-learning? The model-agnostic meta-learning (MAML) [1] is a few shot meta-learning algorithm that uses gradient descent to adapt the model at meta-test time to a new few-shot task, and trains the model parameters at meta-training time to enable rapid adap...

Reading notes: On the Connection Between Adversarial Robustness and Saliency Map Interpretability

Etmann et al. Connection between robustness and interpretability On the Connection Between Adversarial Robustness and Saliency Map Interpretability Advantage and Disadvantages of adversarial training? While this method – like all known approaches of defense – decreases the accuracy of the classifier, it is also successful in increasing the robustness to adversarial attacks Connections between the interpretability of saliency maps and robustness? saliency maps of robustified classifiers tend to be far more interpretable, in that structures in the input image also emerge in the corresponding saliency map How to obtain saliency maps for a non-robustified networks? In order to obtain a semantically meaningful visualization of the network’s classification decision in non-robustified networks, the saliency map has to be aggregated over many different points in the vicinity of the input image. This can be achieved either via averaging saliency maps of noisy versions of the image (Smilkov...

Study notes: Multi-Pointer Co-Attention Networks for Recommendation

Traditional collaborative filtering methods usually only incorporate user-item rating pairs for recommendation, the vast available metadata is just ignored in such scenario. With the recent rapid developments of deep learning techniques, neural based recommendation methods is emerging. Most of them benefit from the metadata that improves personalized recommendation significantly. This  paper  is an example that is based on neural architechture for recommendation with user reviews. In this post, I just explain the model itself, for detail experiments and backgrounds, please refer to the original  paper . Problem Formulation Inputs: User ID $a$, Item ID $b$,  user $a$'s reviews set $\boldsymbol{d}_a= \{d_{a1},..., d_{al_a}\}$ and item $b$'s reviews set $\boldsymbol{d}_b= \{d_{b1},..., d_{bl_b}\}$. Note that  $\boldsymbol{d}_a$ contains all reviews given by user $a$, similarly, $\boldsymbol{d}_b$ contains all reviews received by item $b$. Outputs: ...