Content-Based Recommendation Using Machine Learning
Content-Based Recommendation Using Machine Learning
1
Hefei University of Technology, Hefei, Anhui Province, [Link]
tyf88525319@[Link]
2
University of Science and Technology Beijing, Beijing, [Link]
szy2267246060@[Link]
3
Johns Hopkins University, Baltimore, Maryland, USA
zyao5@[Link]
rized licensed use limited to: SVKM¿s NMIMS Mukesh Patel School of Technology Management & Engineering. Downloaded on December 20,2022 at 05:56:40 UTC from IEEE Xplore. Restrictions
2. RESEARCH METHOD ..., common Words [100] in review. Then, SVM is used for
each of the categories and obtain the category with the highest
The dataset in this work is extracted from reviews written for score as predicted category.
items purchased by Amazon users, where both the users and
the items are anonymized. In the dataset, 200,000 data are
2.3. Rating Prediction
in the form of JSON, which includes features such as pur-
chased item-ID, category, rating, price and user review, etc. The baseline model is to predict only according to the mean
The dataset is split into train set and dev set by the ratio of score of the training examples. And the performance of the
8:2. model is measured by mean square error. Then three kind of
To structure a setup for our model, this work is con- improvement are implemented based on neural networks.
structed in the following approach for each part: First, a base-
line model for each task is provided, which used when there
2.3.1. BERT-Mean
is no profile for a customer (i.e., no purchase history). Then,
a simple algorithm is used for comparison to the baseline Firstly, feed the sentences into the BERT base network [15],
model. Finally, apply a much more sophisticated algorithm and a tensor of size [batch-size, seq-len, hidden-size] can be
proposed by this work for information retrieval. We evaluate obtained. Then mean over the second dimension, which is
all the models on the held-out validation set. The structure mean over the timesteps, get a tensor, size of [batch size, hid-
of the method is shown in figure Fig. 4. The final task is to den size]. Finally, this tensor is fed into a linear regression
generate the user profiles for the recommender system. And layer to get the final output for the rating scores. The archi-
this final task is divided into three parts: Purchase Prediction, tecture of this model is shown in Fig. 1.
Category Prediction and Rating Prediction. For each part of
the task, the baseline model and the improved models are Mean Over Squeeze &
Input BERT Output
provided and compared. Timestep Linear Layer
rized licensed use limited to: SVKM¿s NMIMS Mukesh Patel School of Technology Management & Engineering. Downloaded on December 20,2022 at 05:56:40 UTC from IEEE Xplore. Restrictions
3. RESULTS So, for Jaccard Sim, the performance increases greatly. As for
logistic regression, the popular items feature and Jaccard sim
For task 1, we utilize the accuracy as the metric to measure the feature were combined, so the accuracy is further increased.
performance in purchase prediction. The accuracy is defined As for Category prediction. The baseline model provides
as times of correct prediction divided by the total times of a preliminary prediction when user does not leave a review.
prediction. As is shown is table 1, adding Jaccard Similarity SVM extends upon the baseline to retrieve and analyze infor-
has greatly improved the performance of purchase prediction. mation from text, which is a good improvement.
And logistic regression model further boosts the performance As for Rating Prediction. We noticed that the result for
and achieved the best accuracy is this task. Finally, logistic Bert is worse than simple Conv + LSTM. Since the solo out-
regression outperformed the baseline model by 37.8%. put of Bert may not be a good summary of the whole input’s
Table 1. Purchase Prediction semantic representations, this will result in poor performance
in the downstream tasks. While the Conv and LSTM took
Model Name Accuracy
average over timesteps, which is exactly the way should be
Baseline 0.566 taken to get the precise semantic representations.
Jaccard Sim 0.72 In conclusion, this work can provide a reasonable user
Logistic Regression 0.78 profile for a recommender system.
However, in this work, only text information was used for
As for category prediction, similarly, the measurement recommender system. We would like to further investigate
metric is also the accuracy of the model’s prediction. And about how to recommend with help of video signals. For
the results are shown in table 2, the support vector machine example, if allowed, a model can be developed to analyze
(SVM) model outperformed the baseline model by 35.6%, the face expression captured by a camera for recommenda-
which indicates that SVM can handle this task reasonably. tion. To be specific, some real time facial expression recog-
Table 2. Category Prediction nition techniques [16, 17] can be applied to capture the facial
expression. And some user motion driven recommendation
Model Name Accuracy
techniques [18] can be applied to analyze user emotion from
Baseline 0.57 users’ facial expression. And the recommendation can even
SVM 0.7731 be based on gaze [19] and other interesting features captured
by video camera with the help of speeding up robust features
And as for rating prediction, mean square error (MSE) is extraction algorithms [20].
utilized as the measurement metric. MSE is defined as equa-
tion below. 5. REFERENCES
n
1X
MSE = (Yi − Ŷi )2 (2)
n i=1 [1] Paul Resnick and Hal R Varian, “Recommender sys-
tems,” Communications of the ACM, vol. 40, no. 3, pp.
Where n is the total number of predictions; Yi denotes the
56–58, 1997.
observed value and Ŷi denotes the predicted value.
And note that the lower the MSE is, the better the perfor- [2] Jie Lu, Dianshuang Wu, Mingsong Mao, Wei Wang, and
mance is. As shown in table 3, conv-LSTM got the lowest Guangquan Zhang, “Recommender system application
MSE value and outperformed the baseline model by 63.4%. developments: a survey,” Decision Support Systems,
Thus, this work can achieve a reasonable user profile for the vol. 74, pp. 12–32, 2015.
recommender system.
[3] Gediminas Adomavicius and Alexander Tuzhilin, “To-
Table 3. Rating Prediction ward the next generation of recommender systems: A
Model Name MSE survey of the state-of-the-art and possible extensions,”
Baseline 1.23 IEEE transactions on knowledge and data engineering,
BERT-Mean 0.58 vol. 17, no. 6, pp. 734–749, 2005.
BERT-Two-Layer 0.57 [4] Yoon Ho Cho, Jae Kyeong Kim, and Soung Hie Kim,
Con-LSTM 0.45 “A personalized recommender system based on web us-
age mining and decision tree induction,” Expert systems
with Applications, vol. 23, no. 3, pp. 329–342, 2002.
4. DISCUSSION AND CONCLUSIONS
[5] Soulef Benhamdi, Abdesselam Babouri, and Raja
As for Purchase Prediction. Since people tend to purchase Chiky, “Personalized recommender system for e-
popular item but not necessarily true, so the accuracy around learning environment,” Education and Information
0.5 is reasonable. Similar users also tend to buy similar items. Technologies, vol. 22, no. 4, pp. 1455–1477, 2017.
rized licensed use limited to: SVKM¿s NMIMS Mukesh Patel School of Technology Management & Engineering. Downloaded on December 20,2022 at 05:56:40 UTC from IEEE Xplore. Restrictions
User
Profile
BERT
Baseline Jaccard Baseline Baseline BERT Conv-
Model Model LR Model SVM
Model Mean
Two
LSTM
layer
[6] David Goldberg, David Nichols, Brian M Oki, and Dou- [14] Shristi Shakya Khanal, PWC Prasad, Abeer Alsadoon,
glas Terry, “Using collaborative filtering to weave an in- and Angelika Maag, “A systematic review: machine
formation tapestry,” Communications of the ACM, vol. learning based recommendation systems for e-learning,”
35, no. 12, pp. 61–70, 1992. Education and Information Technologies, vol. 25, no. 4,
pp. 2635–2664, 2020.
[7] Michael J Pazzani and Daniel Billsus, “Content-based
recommendation systems,” in The adaptive web, pp. [15] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and
325–341. Springer, 2007. Kristina Toutanova, “Bert: Pre-training of deep bidirec-
tional transformers for language understanding,” arXiv
[8] Aneesh Sharma, Jerry Jiang, Praveen Bommannavar, preprint arXiv:1810.04805, 2018.
Brian Larson, and Jimmy Lin, “Graphjet: Real-time
content recommendations at twitter,” Proceedings of the [16] Philipp Michel and Rana El Kaliouby, “Real time fa-
VLDB Endowment, vol. 9, no. 13, pp. 1281–1292, 2016. cial expression recognition in video using support vector
machines,” in Proceedings of the 5th international con-
[9] Yi Cai, Ho-fung Leung, Qing Li, Huaqing Min, Jie ference on Multimodal interfaces, 2003, pp. 258–264.
Tang, and Juanzi Li, “Typicality-based collaborative fil-
[17] Huiyuan Yang, Umur Ciftci, and Lijun Yin, “Facial ex-
tering recommendation,” IEEE Transactions on Knowl-
pression recognition by de-expression residue learning,”
edge and Data Engineering, vol. 26, no. 3, pp. 766–779,
in Proceedings of the IEEE conference on computer vi-
2013.
sion and pattern recognition, 2018, pp. 2168–2177.
[10] J Ben Schafer, Dan Frankowski, Jon Herlocker, and Shi-
[18] Mahesh Babu Mariappan, Myunghoon Suk, and Balakr-
lad Sen, “Collaborative filtering recommender systems,”
ishnan Prabhakaran, “Facefetch: A user emotion driven
in The adaptive web, pp. 291–324. Springer, 2007.
multimedia content recommendation system based on
[11] Ján Suchal and Pavol Návrat, “Full text search engine as facial expression recognition,” in 2012 IEEE Interna-
scalable k-nearest neighbor recommendation system,” tional Symposium on Multimedia. IEEE, 2012, pp. 84–
in IFIP International Conference on Artificial Intelli- 87.
gence in Theory and Practice. Springer, 2010, pp. 165– [19] Saurabh Jaiswal, Shubham Virmani, Vishal Sethi, Kan-
173. jar De, and Partha Pratim Roy, “An intelligent recom-
mendation system using gaze and emotion detection,”
[12] Malte Ludewig, Iman Kamehkhosh, Nick Landia, and
Multimedia Tools and Applications, vol. 78, no. 11, pp.
Dietmar Jannach, “Effective nearest-neighbor music
14231–14250, 2019.
recommendations,” in Proceedings of the ACM Recom-
mender Systems Challenge 2018, pp. 1–6. 2018. [20] Rui Wang, Yijie Shi, and Wenming Cao, “Ga-surf: A
new speeded-up robust feature extraction algorithm for
[13] Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay, “Deep multispectral images based on geometric algebra,” Pat-
learning based recommender system: A survey and new tern Recognition Letters, vol. 127, pp. 11–17, 2019.
perspectives,” ACM Computing Surveys (CSUR), vol.
52, no. 1, pp. 1–38, 2019.
rized licensed use limited to: SVKM¿s NMIMS Mukesh Patel School of Technology Management & Engineering. Downloaded on December 20,2022 at 05:56:40 UTC from IEEE Xplore. Restrictions