Assessing customer satisfaction in digital banking based on reviews: methods of sentiment analysis and topic modeling
Басханов А.Р.1 ![]()
1 Финансовый университет при Правительстве Российской Федерации, Москва, Россия
Статья в журнале
Маркетинг и маркетинговые исследования (РИНЦ, ВАК)
опубликовать статью
Том 31, Номер 4 (Октябрь-декабрь 2026)
Introduction
The rapid development of digital financial services has significantly increased the importance of analyzing user experience and customer perceptions of banking services. In recent years, large digital platforms and marketplaces have actively developed their own financial ecosystems by integrating banking products into existing e-commerce services. A notable example of this transformation is Ozon, which has evolved from a traditional marketplace into a platform with its own financial division, Ozon Bank. The introduction of a banking service has allowed the platform to expand its range of services by offering additional financial instruments, including loyalty programs and discounts for purchases made within the ecosystem. Such integration of e-commerce and financial services creates a new user experience and attracts considerable customer attention.
Customer reviews represent an important source of information about service quality and user satisfaction with financial products. Consequently, the large volume of textual data generated on online platforms requires systematic analysis. The study of user feed-back makes it possible to identify both the strengths and weaknesses of financial products and services, thereby supporting more informed decisions regarding service improvement. The analysis of large-scale textual datasets typically relies on sentiment analysis and topic modeling techniques based on machine learning methods [12]. In addition to identifying the sentiment of textual content, these approaches allow researchers to detect clusters representing key discussion topics and to predict customer satisfaction ratings [17].
The aim of this study is to conduct a comprehensive analysis of customer reviews of Ozon Bank published on the online platform Banki.ru. The study employs data parsing, statistical analysis, topic modeling, and multitask neural network classification. To achieve this goal, the following tasks were defined:
· collecting reviews and their attributes through web page parsing;
· analyzing the relationship between the main grade and additional grades;
· examining the geographical and temporal characteristics of the reviews;
· identifying topics discussed in the reviews using BERTopic and clustering methods [10];
· developing a multitask classification model based on RuBERT to estimate missing grades [6];
· evaluating model performance and comparing customer satisfaction assessments after incorporating predicted values into the analysis.
The study has practical relevance, as the model enables a more accurate assessment of user feedback by integrating quantitative ratings with the sentiment of textual data. To facilitate interpretation and reuse, Figure 1 presents a schematic overview of a typical application scenario for the dataset.
Figure 1. Schematic workflow illustrating a typical application scenario Source: created by the author.
Materials and methods
To conduct the sentiment analysis, customer reviews related to Ozon Bank were collected from the online platform Banki.ru [1]. This bank is ranked among the top five institutions in the platform’s “People’s Rating”, indicating a high level of customer engagement and rapidly growing popularity. Consequently, the platform provides a sufficiently large volume of user-generated content suitable for analysis. The final dataset consisted of 22,193 reviews.
Within the scope of this study, the following data attributes were extracted during the parsing process:
· row index column generated in Excel (integer);
· direct link to each review (string);
· username of the reviewer (string);
· user’s location (string);
· review publication date (date);
· review publication time (time);
· overall rating on a five-point scale (integer);
· additional ratings on a three-point scale (integer) according to the following criteria (transparent conditions, polite staff, accessibility and support, app and website usability);
· review text (string).
The overall rating is always available, while missing additional ratings are assigned “0” and excluded from calculations.
Table 1 presents an example of the resulting dataset structure.
Table 1
Dataset structure
|
URL
|
Username
|
City
|
Date
|
Time
|
Grade
|
Transparent
conditions
|
Polite staff
|
Accessibility and
support
|
App and website usability
|
Review
|
|
https://...
|
user-70…
|
Omsk
|
15.01.2026
|
12:15
|
5
|
2
|
3
|
3
|
3
|
recently…
|
|
https://...
|
user-89…
|
Moscow
|
14.01.2026
|
18:17
|
5
|
3
|
3
|
3
|
3
|
good…
|
|
https://...
|
user-17…
|
Ufa
|
14.01.2026
|
14:51
|
5
|
0
|
3
|
0
|
3
|
checked…
|
|
https://...
|
user-61…
|
Bryansk
|
14.01.2026
|
09:25
|
5
|
3
|
3
|
3
|
3
|
using…
|
To identify which criteria receive the highest grades from users, an analysis of the average values of additional ratings depending on the overall rating was conducted. Figure 2 presents the distribution of average additional ratings by overall score, where higher values are highlighted using a horizontal gradient, and the rightmost column shows the overall average of additional criteria for each rating. Regardless of the overall rating, users consistently rate “Polite staff” and “App and website usability” highly, making these key service strengths.
Figure 2. Average ratings across additional criteria Source: created by the author.
Since it was previously noted that users on Banki.ru may omit some additional ratings when submitting reviews, it became necessary to analyze not only the average values but also the frequency of ratings assigned to each criterion. Figure 3 presents the distribution of rating frequencies across the additional criteria.
Figure 3. Number of reviews across additional criteria Source: created by the author.
The rightmost column represents the total number of overall ratings. It can be observed that, although users do not always evaluate all additional criteria, the “Accessibility and support” criterion is reported most frequently. This may indicate that this criterion is either the most important for users or the most intuitive.
The dataset also includes timestamps, allowing the construction of an hourly activity distribution. Figure 4 shows review counts, highlighting peak periods. In addition to the temporal distribution analysis, the distribution of reviews by year was examined. Figure 5 presents the annual distribution of reviews.
Figure 4. Distribution of reviews by hour Source: created by the author.
Figure 5. Distribution of reviews by year Source: created by the author.
The analysis of the monthly distribution indicates that the growth in the number of reviews in 2025 was relatively uniform throughout the year, suggesting that there was likely no artificial inflation of reviews. The probable factors contributing to this increase include:
· the launch of lending services for both individual and corporate customers;
· improvements to existing financial products;
· introduction of a new customer loyalty program;
· growth of the customer base;
· expansion of the bank’s geographical presence.
These changes likely enhanced user engagement and encouraged more reviews on the platform Banki.ru.
A map using city coordinates from an open-source GitHub repository was created to illustrate the geographical distribution of reviews. Figure 6 shows this spatial distribution. The results indicate that the majority of reviews were submitted by users from the European part of Russia. This suggests a high concentration of bank customer activity in the central and western regions, likely related to population density and the accessibility of banking services in these areas.
Figure 6. Geographical distribution of reviews Source: created by the author.
When analyzing reviews, it is also important to assess their authenticity. Patterns of similar reviews were identified. In Table 2, text highlighted in black represents the identical portions, while red indicates the differing segments. Since only “Verified” reviews were collected, such a high degree of similarity was unexpected; in some cases, the only change was the grammatical gender in the text. The reviews are presented in lowercase letters and without punctuation, as the data were standardized during the preprocessing stage [16].
Table 2
Examples of similar reviews
(original authors’ wording preserved in Russian reviews)
|
Pattern
|
Review (Russian Original)
|
Review (English Translation)
|
|
1
|
заморозили
счёт в банке написал в поддержку сутки не отвечают мне в чате…
|
My bank account was frozen. I contacted
support, but they haven't replied in the chat for 24 hours…
|
|
ограничили счёт в банке написал в поддержку
уже вторые
сутки не отвечают
мне в чате…
|
My bank
account was restricted. I contacted support, but they still haven't
replied in the chat for over 48 hours…
| |
|
2
|
возникла
проблема заказала товар с постоплатой но оказалось
что арест на озон карте проверила
на госуслугах номер постановления…
|
I
ran into a problem. I ordered (female) an item with post-payment, but it turned out that
my Ozon Card was under seizure. I checked (female) the enforcement order
number on the government services portal…
|
|
возникла проблема заказал товар с постоплатой но оказалось
что арест на озон карте проверил
на госуслугах номер постановления…
|
I ran into a problem.
I ordered (male)
an item with post-payment, but it turned out that my Ozon Card was under
seizure. I checked (male)
the enforcement order number on the government services portal…
|
It can be assumed that, during the verification process, a review is not compared against previously submitted reviews. This may allow banks to artificially inflate their ratings with positive reviews or enable competitors to post negative reviews. To mitigate such issues, suspicious reviews should either be periodically removed or immediately checked for uniqueness.
The platform Banki.ru is particularly suitable for research purposes because each review contains a wide range of associated attributes, including ratings and additional information. Moreover, the review search interface includes a built-in filtering system that facilitates the formation of a relevant dataset for analysis. Figure 7 illustrates the review filtering interface on the platform Banki.ru.
Figure 7. Review filtering system Source: screenshot from Banki.ru [1].
In the “Services” section, the following categories are available to users:
· all;
· individual customers, including debit card, credit card, mortgage, car loan, consumer loan, restructuring/refinancing, deposit, money transfer, remote banking services for individuals, other (individual customers), mobile application, customer service for individuals;
· corporate customers, including settlement and cash services, acquiring services, pay-roll project, deposit, business lending, bank guarantee, leasing, other services, remote banking services for legal entities, business mobile application, customer service for legal entities.
Since the topics of the reviews are identified using the BERTopic model, the “All” option was selected in the “Services” section in order to avoid preliminary thematic filtering.
In the “Reviews” section, the following options are available:
· all reviews;
· with the bank response;
· with the resolved issue;
· verified.
To prevent manipulated or artificially generated reviews from entering the dataset, the “Verified” filter option was selected. It should be noted that each review on the platform has one of three statuses: “not accepted”, “under verification”, or “accepted”. As a result, instead of the total 45,049 available reviews, only 22,193 reviews satisfied the specified filtering conditions. This reduction occurs because approximately half of the reviews have the status “not accepted”, while the most recent reviews (typically from the last week) remain “under verification”.
In the “Grade” section, the following options are available:
· any grade;
· positive reviews;
· negative reviews.
To predict missing ratings, the dataset included both positive and negative reviews; therefore, the “Any grade” option was selected.
Figure 8 illustrates the rating system based on additional criteria. It also displays supplementary information, including the reviewer’s username, user location, review publication date and time, review text, as well as the overall rating.
Figure 8. Review rating system on the platform Source: screenshot from Banki.ru [1].
To ensure the stability of the data collection process, several measures were implemented in the parsing script when interacting with the platform Banki.ru:
· random time intervals were introduced between consecutive actions to simulate human behavior and reduce the likelihood of automated traffic detection;
· different User-Agent headers were configured within browser sessions;
· a pool of authenticated proxy servers was utilized. In the event of an HTTP 403 error, the proxy was automatically rotated.
The following sections describe the procedures used to collect, clean, and structure the data, including the extraction of additional criteria ratings, the preprocessing of review texts, and the preparation of geographic information.
In Figure 8, data collection for the additional criteria was carried out as follows. First, we checked which criteria were present for each review; if a criterion was absent, it was assigned a score of 0, and these zero-value entries were not included in the subsequent analysis. For the criteria that were present, the rating was extracted based on the number of blue-filled segments out of three, since the score was not explicitly provided as a numeric value. The review text was also included in the dataset, with punctuation removed and all characters converted to lowercase to facilitate text analysis [9]. This preprocessing improves BERTopic topic modeling performance [10].
To generate Figure 6, which illustrates the geographic distribution of reviews, we used Python with the Pandas and GeoPandas libraries for data processing and Matplotlib for visualization [8]. To obtain geographic coordinates, we used an open GitHub dataset containing coordinates of Russian cities [13]. In our dataset, the “City” column sometimes contained only the city name (e.g., “Ufa”) and sometimes included additional information in parentheses (e.g., “Ufa (Republic of Bashkortostan)”). These entries were not removed from the dataset; instead, the text within parentheses was ignored when matching cities with their coordinates for the map. Review counts were then aggregated by city and plotted to visualize the spatial distribution of reviews.
Results
After completing the analysis of the quantitative characteristics of the reviews, RuBERT, a transformer-based model for Russian text analysis, was employed to evaluate sentiment in all 22,193 reviews [15]. Each review was preprocessed by converting text to lowercase and removing punctuation to improve model performance [3]. The model assigned each review to one of three sentiment classes: negative, neutral, or positive [11]. The distribution of predicted sentiments was then calculated and visualized to provide an overview of the dataset’s overall sentiment. Figure 9 presents the results of the sentiment analysis.
Figure 9. Sentiment analysis of reviews Source: created by the author.
It is important to recall the data from Figure 3, which shows that 14,896 reviews received an overall rating of “5”, representing nearly 70% of all reviews. However, RuBERT classified the majority of these highly positive reviews as “neutral”. From a customer experience perspective, the sentiment analysis suggests that the overall ratings alone do not fully reflect users’ perceptions, highlighting the nuances that numerical scores may over-look.
Thus, user ratings and sentiment analysis results do not always align [4]. For example, a customer may report a serious issue but still give a “5” grade if the problem was resolved quickly. However, this does not negate the fact that the bank faces challenges that create difficulties for users. Therefore, combining quantitative ratings with textual analysis is necessary to obtain a more accurate interpretation of reviews.
Next, using BERTopic – an approach that has proven effective for identifying thematic structures in unstructured text data across diverse domains [5] – common topics appearing in both positive and negative reviews were identified to explicitly highlight the bank's strengths and weaknesses as perceived by users. These topics were grouped based on keywords, enabling the identification of main areas discussed in reviews, ranging from card services and cashback to customer support and service-related issues. The results of this analysis are presented in Table 3, which shows the keywords associated with each topic and the proportion of reviews of each sentiment corresponding to that topic.
The most prominent topic among negative reviews is “Customer service and sup-port”, which aligns with the observations from Figure 2. If the “Accessibility and support” criterion is considered equivalent, its average rating is 1.04, confirming that customer dissatisfaction is primarily associated with service and support issues.
Conversely, among positive reviews, the most common topics are savings accounts, cashback and bonuses, cards and transactions, and support and refunds.
Table 3
Topics mentioned in reviews
|
Topic
|
Keywords
|
Share (%)
|
|
Negative topics (1,448 reviews)
| ||
|
Customer service and support
|
employee, operator, question, answer
|
27,6%
|
|
Stars and bonuses
|
stars, ruble, purchases, loyalty
|
25,0%
|
|
Cashback and categories
|
cashback, cashback rewards, purchases, categories
|
24,9%
|
|
Cards and virtual cards
|
card, plastic, virtual
|
22,5%
|
|
Positive topics (10,424 reviews)
| ||
|
Savings accounts
|
savings, account, rate, interest
|
10,6%
|
|
Cashback and categories
|
cashback, cashback rewards, purchases, categories
|
9,1%
|
|
Cards and transactions
|
card, virtual, plastic, transfer
|
8,6%
|
|
Support and refunds
|
operator, refund, thank you, contacted
|
8,2%
|
|
Documents and identification
|
passport, data, photo
|
6,7%
|
|
Limits and installments
|
limit, installment, loan, account
|
6,2%
|
|
Notifications and subscriptions
|
notifications, subscription, premium, return
|
5,8%
|
|
ATMs and cash withdrawals
|
ATM, fee, cash, withdrawal
|
5,3%
|
|
Other operations
|
money, blocked, account, funds
|
4,8%
|
|
Automation and chatbot
|
robot, responds, operator, question
|
3,8%
|
Before developing a model to predict missing ratings, it is necessary to assess the extent of missing data. The number of missing ratings for each additional criterion was calculated. The results are presented in Table 4, which highlights the aspects for which users most frequently did not provide ratings.
Table 4
Proportion of grades for each criterion
|
Criterion
|
Grade
| |||
|
No grade
|
1
|
2
|
3
| |
|
Transparent
conditions
|
1 397
|
4 471
|
2 304
|
14 021
|
|
Polite
staff
|
1 233
|
3 575
|
1 771
|
15 614
|
|
Accessibility
and support
|
1 019
|
5 616
|
550
|
15 008
|
|
App
and website usability
|
1 247
|
3 593
|
2 226
|
15 127
|
By leveraging user ratings and the RuBERT model, a multitask classification model can be implemented to predict missing additional ratings on a scale from 1 to 3 based on the review text.
In this approach, each review is simultaneously analyzed across four criteria [6]. This enables the reconstruction of inferred user ratings in cases where explicit ratings were not provided, but the review text clearly reflects the customer’s sentiment toward the corresponding aspect of service.
Model performance was evaluated using accuracy and Macro-F1 metrics on the original dataset, without incorporating the newly predicted values. Accuracy was calculated according to Equation (1), while the Macro-F1 score was computed using Equation (2).
|
|
|
|
where K is the number of classes, and
and
denote precision and recall for class k, respectively.
The results are presented in Table 5.
Table 5
Model evaluation
|
Criterion
|
Accuracy
|
Macro-F1
|
|
Transparent conditions
|
0.887
|
0.695
|
|
Polite staff
|
0.917
|
0.753
|
|
Accessibility and support
|
0.972
|
0.724
|
|
App and website usability
|
0.875
|
0.664
|
The extremely high accuracy can be explained by the fact that approximately 75% of reviews in the training dataset had a rating of “3”. This class imbalance encourages the model to predict the most frequent rating. In other words, when one class dominates, it is statistically difficult for the model to make a mistake, which artificially inflates the accuracy metric.
At the same time, the Macro-F1 score (ranging from 0.66 to 0.75) provides a more realistic assessment of model performance [15]. Because Macro-F1 averages the F1 score across all classes regardless of their frequency, it is sensitive to the model’s performance on underrepresented classes (grades “1” and “2”). Lower Macro-F1 values indicate that the model struggles to predict these less frequent ratings, which is due to their limited representation in the training data. For example, for the criterion “Accessibility and support”, a grade of “2” appears only 550 times, making the prediction of this class unlikely [7].
Thus, the high accuracy reflects the imbalance in the training dataset rather than incorrect model settings. The dataset itself is inherently imbalanced, as users predominantly leave positive reviews for Ozon Bank.
To visually demonstrate the impact of the predicted values, Figure 10 was constructed based on the data from Figure 2. It shows the change in average ratings for each additional criterion after filling in missing values using the model. Only changes with an absolute value greater than 0.01 were displayed to exclude statistically insignificant fluctuations and focus on meaningful shifts.
Figure 10. Average ratings for additional criteria after applying the model Source: created by the author.
The results obtained after applying the model demonstrate that even a small number of missing ratings can affect the overall average values. By combining sentiment analysis with the multitask model, a more accurate assessment of user feedback is achieved.
Discussion
The phenomenon of Ozon Bank’s rapid rise in popularity in the Russian financial market provides an interesting case for analyzing the success of a business model that integrates a traditional marketplace with banking services [14]. This explosive growth has provoked strong dissatisfaction among the largest traditional Russian banks, which even appealed to the Government of the Russian Federation to introduce certain restrictions on Ozon Bank. This discontent was primarily driven by the fact that Ozon Bank offers significant discounts on purchases made through the Ozon marketplace, which clearly constitutes a strong incentive for consumers to use the bank’s financial services.
At present, the majority of Ozon marketplace clients use Ozon Bank’s services mainly to make discounted purchases on the platform. However, Ozon Bank itself is evidently interested in evolving into a fully-fledged bank, rather than merely serving as a tool for shopping. Thus, based on an analysis of user experience, this study allows for evaluating clients’ preferences and needs in using Ozon Bank as a conventional banking service.
Conclusion
In this study, a total of 22,193 reviews about Ozon Bank were extracted from the platform Banki.ru. The dataset was structured and included additional attributes such as temporal and geographic information. Statistical analysis showed that users consistently value staff courtesy and application and website usability, regardless of their overall impression of the bank. The most frequently rated criterion was “Accessibility and support”, highlighting its importance in evaluating and perceiving service quality.
Topic modeling using BERTopic identified the main themes mentioned in user reviews. Negative reviews primarily focused on service and support issues, bonus accrual, and card operations, while positive reviews were mostly related to savings accounts, cashback, card services, and issue resolution. Comparison of the original quantitative ratings with sentiment-based assessments revealed substantial differences, emphasizing the need to combine quantitative and qualitative analyses when interpreting customer experience.
A multitask classification model based on RuBERT was developed to predict missing ratings in reviews. Despite high accuracy scores, Macro-F1 analysis highlighted the impact of class imbalance and the limited predictive performance for less frequent ratings. The model allowed adjustment of average criterion scores, providing a more accurate representation of user perception. Overall, the integration of web scraping, statistical analysis, topic modeling, and neural network classification demonstrates high effectiveness in studying customer reviews. The findings can be applied to improve customer service, enhance the quality of banking products, and develop systems for monitoring user satisfaction [2].
Страница обновлена: 31.08.2026 в 12:52:38
Assessing customer satisfaction in digital banking based on reviews: methods of sentiment analysis and topic modeling
Baskhanov A.R.Journal paper
Marketing and marketing research
Volume 31, Number 4 (October-December 2026)
