Альтернативный подход к отраслевой классификации публичных компаний на основе текстовой кластеризации: эмпирическое исследование на данных NASDAQ
Ветрова М.А.1
, Купоров В.С.2
, Горкавцев М.О.3 ![]()
1 Санкт-Петербургский государственный университет, Санкт-Петербург, Россия
2 VERSUS, Санкт-Петербург,, Россия
3 Технологии Доверия - Консультирование, Санкт-Петербург, Россия
Статья в журнале
Информатизация в цифровой экономике (РИНЦ, ВАК)
опубликовать статью | оформить подписку
Том 7, Номер 2 (Апрель-июнь 2026)
Аннотация:
Экспертные отраслевые классификаторы (GICS на международных рынках, ОКВЭД в Российской Федерации) лежат в основе большинства эмпирических исследований в финансах и экономике, однако плохо отражают гибридные бизнес-модели, возникающие в результате цифровой трансформации. Цель исследования – разработать и апробировать воспроизводимый индуктивный подход к группировке компаний на основе их текстовых самоописаний, оптимизирующий соотношение качества кластеризации и вычислительной экономичности. На выборке 3 380 публичных компаний, торгующихся на бирже NASDAQ, применена методологическая связка: эмбеддинги Sentence-BERT (модель all-MiniLM-L6-v2) без дообучения, снижение размерности методом главных компонент и кластеризация алгоритмом k-средних. Устойчивость результатов подтверждена тремя независимыми способами: двумя альтернативными конфигурациями (без снижения размерности и через плотностную кластеризацию UMAP+HDBSCAN), формальными метриками согласия с GICS (скорректированный индекс Рэнда 0,46, средняя чистота 61,4%) и финансовым профилем кластеров по семи показателям, не участвовавшим в построении кластеров. Получены шесть содержательно интерпретируемых групп: два кластера (клинические биотехнологии и коммерческие банки) воспроизводят отраслевую структуру с чистотой 97-98%, два других выявляют структурные феномены, неразличимые в GICS (компании-оболочки в рамках сделок слияния и поглощения и потребительский гибрид, объединяющий ритейл с B2C-платформами). Предложенная методология обладает высокой воспроизводимостью, не требует дообучения модели и применима для уточнения отраслевой структуры на данных российских компаний с использованием ОКВЭД в качестве референсной классификации.
Ключевые слова: отраслевая классификация, текстовая кластеризация, Sentence-BERT, метод k-средних, GICS, цифровая трансформация
JEL-классификация: G14, C38, L25
Источники:
2. Козырь Н.С., Коваленко В.С. Метрика отраслевой классификации в Российской Федерации и за рубежом // Экономический анализ: теория и практика. – 2017. – № 10(469). – c. 1914-1927. – doi: 10.24891/ea.16.10.1914.
3. Макеева Е.Ю., Аршавский И.В. Применение нейронных сетей и семантического анализа для прогнозирования банкротства // Корпоративные финансы. – 2014. – № 4(32). – c. 130-141. – doi: 10.17323/j.jcfr.2073-0438.8.4.2014.130-141.
4. Михненко П.А. Трансформация деловой лексики годовых отчетов крупнейших российских компаний: Data Mining // Управленец. – 2022. – № 5. – c. 17-33. – doi: 10.29141/2218-5003-2022-13-5-2.
5. Федорова Е.А., Сальникова П.А. Влияние раскрытия информации об экологических инициативах на цены акций публичных компаний России // Экономический журнал. – 2024. – № 2. – c. 223-247. – doi: 10.17323/1813-8691-2024-28-2-223-247.
6. Bai H., Xing F. Z., Cambria E., Huang W.-B. Business Taxonomy Construction Using Concept-Level Hierarchical Clustering // Proceedings of the First Workshop on Financial Technology and Natural Language Processing. Macao, China, 2019. – p. 1-7.– doi: 10.48550/arXiv.1906.09694.
7. Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Minneapolis, Minnesota, 2019. – p. 4171-4186.– doi: 10.18653/v1/N19-1423.
8. Hoberg G., Phillips G. Product Market Synergies and Competition in Mergers and Acquisitions: A Text-Based Analysis // Review of Financial Studies. – 2010. – № 10. – p. 3773-3811. – doi: 10.1093/rfs/hhq053.
9. Hoberg G., Phillips G. Text-Based Network Industries and Endogenous Product Differentiation // Journal of Political Economy. – 2016. – № 5. – p. 1423-1465. – doi: 10.1086/688176.
10. Jagrič T., Herman A. AI Model for Industry Classification Based on Website Data // Information (Switzerland). – 2024. – № 2. – p. 89. – doi: 10.3390/info15020089.
11. Kim D., Kang H.G., Bae K., Jeon S. An artificial intelligence-enabled industry classification and its interpretation // Internet Research: Electronic Networking Applications and Policy. – 2022. – № 2. – p. 406-424. – doi: 10.1108/INTR-05-2020-0299.
12. McInnes L., Healy J., Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. ArXiv preprint. [Электронный ресурс]. URL: https://arxiv.org/abs/1802.03426.
13. Ortakci Ya., Borhan B. Optimizing SBERT for long text clustering: two novel approaches with empirical insights // The Journal of Supercomputing. – 2025. – № 8. – p. 950. – doi: 10.1007/s11227-025-07414-4.
14. Papenkov M., Meredith C., Noel C., Padalkar J., Hendrickson T., Nitiutomo D., Farrell T. Multi-Industry Simplex: A Probabilistic Extension of GICS. arXiv preprint. [Электронный ресурс]. URL: https://arxiv.org/abs/2310.04280.
15. Reimers N., Gurevych I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. Hong Kong, China, 2019. – p. 3982-3992.– doi: 10.18653/v1/D19-1410.
16. Vamvourellis D., Tóth M., Bhagat S., Desai D., Mehta D., Pasquali S. Company Similarity using Large Language Models. ArXiv preprint. [Электронный ресурс]. URL: https://arxiv.org/abs/2308.08031.
17. Financial statement data for publicly traded companies. Yahoo Finance. [Электронный ресурс]. URL: https://finance.yahoo.com/ (дата обращения: 16.05.2026).
18. Yang H., Lee H. J., Cho S., Cho E. Automatic Classification of Securities using Hierarchical Clustering of the 10-Ks // 2016 IEEE International Conference on Big Data (Big Data). Washington, DC, 2016. – p. 3936-3943.– doi: 10.1109/BigData.2016.7841069.
Страница обновлена: 24.07.2026 в 19:55:49
An alternative approach to the industry classification of public companies based on text clustering: an empirical study based on NASDAQ data
Vetrova M.A., Kuporov V.S., Gorkavtsev M.O.Journal paper
Informatization in the Digital Economy
Volume 7, Number 2 (April-June 2026)
Abstract:
Expert industry classifiers (GICS in international markets, the All-Russian Classifier of Economic Activities in the Russian Federation) underlie most empirical research in finance and economics, but they poorly reflect hybrid business models emerging as a result of digital transformation. The article aims to develop and test a reproducible inductive approach to grouping companies based on their textual self–descriptions, optimizing the ratio of clustering quality and computational efficiency. A methodological combination was applied to a sample of 3,380 public companies traded on the NASDAQ stock exchange: Sentence-BERT embeddings (all-MiniLM-L6-v2 model) without additional training, dimensionality reduction using the principal component method and clustering based on the k-means algorithm. The stability of the results was confirmed in three independent ways: by two alternative configurations (without dimensionality reduction and through density clustering UMAP+HDBSCAN), formal metrics of agreement with GICS (adjusted Rand index 0.46, average purity 61.4%) and the financial profile of clusters according to seven indicators that were not involved in the cluster construction.
Six meaningfully interpreted groups were obtained: two clusters (clinical biotechnologies and commercial banks) reproduce the industry structure with a purity of 97-98%, the other two reveal structural phenomena that are indistinguishable in GICS (shell companies in mergers and acquisitions and a consumer hybrid combining retail with B2C platforms). The proposed methodology is highly reproducible. It does not require additional model training, and it is applicable to refine the industry structure based on data from Russian companies using the All-Russian Classifier of Economic Activities as a reference classification.
Keywords: industry classification, text clustering, Sentence-BERT, k-means method, GICS, digital transformation
JEL-classification: G14, C38, L25
References:
Bai H., Xing F. Z., Cambria E., Huang W.-B. (2019). Business Taxonomy Construction Using Concept-Level Hierarchical Clustering Proceedings of the First Workshop on Financial Technology and Natural Language Processing. 1-7. doi: 10.48550/arXiv.1906.09694.
Devlin J., Chang M.-W., Lee K., Toutanova K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171-4186. doi: 10.18653/v1/N19-1423.
Fedorova E.A., Salnikova P.A. (2024). Impact of Environmental Initiative Disclosures on Share Prices of Russian Public Companies. Ekonomicheskiy zhurnal. 28 (2). 223-247. doi: 10.17323/1813-8691-2024-28-2-223-247.
Financial statement data for publicly traded companiesYahoo Finance. Retrieved May 16, 2026, from https://finance.yahoo.com/
Gorda A.S. (2025). Digital Transformation of Business Models of Enterprises in the Retail Sector. Omskiy nauchnyy vestnik. Seriya Obschestvo. Istoriya. Sovremennost. 10 (4). 120-126. doi: 10.25206/2542-0488-2025-10-4-120-126.
Hoberg G., Phillips G. (2010). Product Market Synergies and Competition in Mergers and Acquisitions: A Text-Based Analysis Review of Financial Studies. 23 (10). 3773-3811. doi: 10.1093/rfs/hhq053.
Hoberg G., Phillips G. (2016). Text-Based Network Industries and Endogenous Product Differentiation Journal of Political Economy. 124 (5). 1423-1465. doi: 10.1086/688176.
Jagrič T., Herman A. (2024). AI Model for Industry Classification Based on Website Data Information (Switzerland). 15 (2). 89. doi: 10.3390/info15020089.
Kim D., Kang H.G., Bae K., Jeon S. (2022). An artificial intelligence-enabled industry classification and its interpretation Internet Research: Electronic Networking Applications and Policy. 32 (2). 406-424. doi: 10.1108/INTR-05-2020-0299.
Kozyr N.S., Kovalenko V.S. (2017). Performance Metrics of Industrial Classification in the Russian Federation and Abroad. Ekonomicheskiy analiz: teoriya i praktika. 16 (10(469)). 1914-1927. doi: 10.24891/ea.16.10.1914.
Makeeva E.Yu., Arshavskiy I.V. (2014). Integration of Neural Networks and Semantic Interpretation for Bankruptcy Prediction. Korporativnye finansy. 8 (4(32)). 130-141. doi: 10.17323/j.jcfr.2073-0438.8.4.2014.130-141.
McInnes L., Healy J., Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension ReductionArXiv preprint. Retrieved from https://arxiv.org/abs/1802.03426
Mikhnenko P.A. (2022). Transformation of the Largest Russian Companies’ Business Vocabulary in Annual Reports: Data Mining. Upravlenets. 13 (5). 17-33. doi: 10.29141/2218-5003-2022-13-5-2.
Ortakci Ya., Borhan B. (2025). Optimizing SBERT for long text clustering: two novel approaches with empirical insights The Journal of Supercomputing. 81 (8). 950. doi: 10.1007/s11227-025-07414-4.
Papenkov M., Meredith C., Noel C., Padalkar J., Hendrickson T., Nitiutomo D., Farrell T. Multi-Industry Simplex: A Probabilistic Extension of GICSarXiv preprint.. Retrieved from https://arxiv.org/abs/2310.04280
Reimers N., Gurevych I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. 3982-3992. doi: 10.18653/v1/D19-1410.
Vamvourellis D., Tóth M., Bhagat S., Desai D., Mehta D., Pasquali S. Company Similarity using Large Language ModelsArXiv preprint. Retrieved from https://arxiv.org/abs/2308.08031
Yang H., Lee H. J., Cho S., Cho E. (2016). Automatic Classification of Securities using Hierarchical Clustering of the 10-Ks 2016 IEEE International Conference on Big Data (Big Data). 3936-3943. doi: 10.1109/BigData.2016.7841069.
