<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">klj</journal-id><journal-title-group><journal-title xml:lang="ru">Казанский лингвистический журнал</journal-title><trans-title-group xml:lang="en"><trans-title>Kazan linguistic journal</trans-title></trans-title-group></journal-title-group><issn pub-type="epub">3033-8751</issn><publisher><publisher-name>Казанский (Приволжский) федеральный университет</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26907/2658-3321.2023.6.3.388-396</article-id><article-id custom-type="elpub" pub-id-type="custom">klj-53</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ФИЛОЛОГИЯ. ТЕОРЕТИЧЕСКАЯ, ПРИКЛАДНАЯ И СРАВНИТЕЛЬНО-СОПОСТАВИТЕЛЬНАЯ ЛИНГВИСТИКА</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>PHILOLOGICAL STUDIES. THEORETICAL, APPLIED AND COMPARATIVE LINGUISTICS</subject></subj-group></article-categories><title-group><article-title>Методы извлечения терминов в научных текстах (на материале статей по направлению науки о земле)</article-title><trans-title-group xml:lang="en"><trans-title>Methods for Terminology Extraction in Scientific Texts (Based on Articles of Earth Sciences)</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-2603-6242</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Падерина</surname><given-names>T. С.</given-names></name><name name-style="western" xml:lang="en"><surname>Paderina</surname><given-names>T. S.</given-names></name></name-alternatives><bio xml:lang="ru"><p>Падерина Татьяна Сергеевна – Младший научный сотрудник </p><p>Иркутск</p></bio><bio xml:lang="en"><p>Paderina Tatiana Sergeevna – Junior Researcher</p><p>Irkutsk</p></bio><email xlink:type="simple">jana-pad@mail.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Иркутский научный центр Сибирского отделения Российской академии наук</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Irkutsk Scientific Center of Siberian Branch of Russian Academy of Sciences</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2023</year></pub-date><pub-date pub-type="epub"><day>08</day><month>12</month><year>2025</year></pub-date><volume>6</volume><issue>3</issue><fpage>388</fpage><lpage>396</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Падерина T.С., 2025</copyright-statement><copyright-year>2025</copyright-year><copyright-holder xml:lang="ru">Падерина T.С.</copyright-holder><copyright-holder xml:lang="en">Paderina T.S.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://www.kljournal.ru/jour/article/view/53">https://www.kljournal.ru/jour/article/view/53</self-uri><abstract><p>Статья посвящена описанию теоретических и прикладных положений первоначального этапа работы по автоматическому извлечению терминов из научных текстов. Данный этап работы является частью государственного задания научной лаборатории лингво-педагогических исследований по теме «Лингвосемиотическая гетерогенность научной картины мира: теоретическое и лингводидактическое описание». Цель исследования заключается в извлечении терминов из подготовленного корпуса научных текстов, относящихся к определенной предметной области. Для этого был использован корпус научных текстов по направлению Науки о Земле, подготовленный методом случайной выборки при помощи приложения Semantic Scholar. Извлечение терминов при помощи автоматической обработки текстов (АОТ) является перспективным направлением исследования, так как позволяет упростить процесс создания терминосистем или составления онтологии для узкоспециализированных предметных областей. В условиях быстро меняющегося потока информации данный вид работы с текстами, безусловно остается актуальным направлением и позволяет быстрее и эффективнее обрабатывать большие объемы материалов. Однако, необходимо отметить, что автоматическое извлечение терминов (АОТ) не всегда является точным и может содержать ошибки. Поэтому, важно проводить дополнительную проверку и корректировку полученных результатов. Перспективы исследования связаны с совершенствованием существующих инструментов автоматической обработки текстов (АОТ). Кроме этого, анализ извлеченных терминов позволил нам сформировать основу для дальнейших практических исследований по созданию цифрового продукта (цифровой модели определенных терминосистем) для хранения, систематизации и использования терминосистем по определённой узкоспециализированной предметной области.</p></abstract><trans-abstract xml:lang="en"><p>The article describes the theoretical and applied provisions of the initial stage of work on automatic extraction of terms from scientific texts. This stage of the work is a part of the state assignment of the Scientific Laboratory of Linguistic and Pedagogical Research on "Linguosemiotic heterogeneity of scientific picture of the world: theoretical and linguodidactic description". The aim of the research is to extract terms from a prepared corpus of scientific texts relating to a particular subject area. For this purpose, a corpus of scientific texts in the field of Earth Sciences, prepared by random sampling using the Semantic Scholar application, was used. The term extraction by automatic text processing (ATP) is a promising area of research as it simplifies the process of creating terminology systems or ontologies for highly specialized subject areas. With the rapidly changing flow of information, this type of work with texts is undoubtedly still relevant and allows for faster and more efficient processing of large volumes of material. However, it should be noted that automatic term extraction is not always accurate and may contain some errors. Therefore, it is important to carry out additional verification and correction of the results obtained. Prospects for the study are related to the improvement of existing automatic text processing tools. In addition, the analysis of the extracted terms has enabled us to form the basis for further practical research into the creation of a digital product (a digital model of certain terminology systems) for the storage, systematization and use of terminology systems for a certain highly specialized subject area.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>терминология</kwd><kwd>извлечение терминов</kwd><kwd>тематическое моделирование</kwd><kwd>научная коммуникация</kwd></kwd-group><kwd-group xml:lang="en"><kwd>terminology</kwd><kwd>terminology extraction</kwd><kwd>thematic modeling</kwd><kwd>scientific communication</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Дементьева Я.Ю., Бручес Е.П., Батура Т.В. Извлечение терминов из текстов научных статей. Программные продукты и системы/Software &amp; Systems. 2022;35(4):689–697. DOI: 10.15827/0236-235X.140.689-697</mixed-citation><mixed-citation xml:lang="en">Dement`eva Ya.Yu., Bruches E.P., Batura T.V. Terms extraction from texts of scientific papers. Programmny`e produkty` i sistemy`/Software &amp; Systems. 2022;35(4):689–697. DOI: 10.15827/0236-235X.140.689-697 (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Большакова Е.И., Семак В.В. Комбинирование методов для извлечения терминов из научно-технического текста. Интеллектуальные системы. Теория и приложения. 2021;25(4):239–242.</mixed-citation><mixed-citation xml:lang="en">Bol`shakova E.I., Semak V.V. Combining methods to extract terms from scientific and technical text. Intellektual`ny`e sistemy`. Teoriya i prilozheniya. 2021;25(4):239–242. (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Grishman R. Information Extraction. In: The Handbook of Computational Linguistics and Natural Language Processing. A. Clark, C. Fox, and S. Lappin (Eds). WileyBlackwell; 2010. Pp. 515–530.</mixed-citation><mixed-citation xml:lang="en">Grishman R. Information Extraction. The Handbook of Computational Linguistics and Natural Language Processing. A. Clark, C. Fox, and S. Lappin (Eds). WileyBlackwell; 2010. Pp. 515–530.</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Бручес Е. П., Батура Т. В. Метод автоматического извлечения терминов из научных статей на основе слабо контролируемого обучения. Вестник НГУ. Серия: Информационные технологии. 2021;19(2):5–16. DOI 10.25205/1818-7900-2021-19-2-5-16</mixed-citation><mixed-citation xml:lang="en">Bruches E. P., Batura T. V. Method for Automatic Term Extraction from Scientific Articles Based on Weak Supervision. Vestnik NGU. Seriya: Informacionny`e texnologii. 2021;19(2):5–16. DOI 10.25205/1818-7900-2021-19-2-5-16 (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Рогачева В. Э. Методы извлечения терминологических единиц из корпуса сопоставимых текстов. Вестник Воронежского государственного университета. Серия: Лингвистика и межкультурная коммуникация. 2017;(2):118–122.</mixed-citation><mixed-citation xml:lang="en">Rogacheva, V. E`. Methods of extracting terminological units from the corpus of comparable texts. Vestnik Voronezhskogo gosudarstvennogo universiteta. Seriya: Lingvistika i mezhkul`turnaya kommunikaciya. 2017;(2):118–122. (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Eckart de Castilho R., Mújdricza-Maydt, É.,et al. A Web-based Tool for the Integrated Annotation of Semantic and Syntactic Structures. In Proceedings of the LT4DH workshop at COLING. 2016. Osaka, Japan.</mixed-citation><mixed-citation xml:lang="en">Eckart de Castilho R., Mújdricza-Maydt, É.,et al. A Web-based Tool for the Integrated Annotation of Semantic and Syntactic Structures. In Proceedings of the LT4DH workshop at COLING. 2016. Osaka, Japan (In Eng.)</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Шейко А.М. Инструменты прикладной лингвистики в контроле качества перевода. Казанский лингвистический журнал. 2023;6(2):282–293. DOI 10.26907/2658-3321.2023.6.2.282-293.</mixed-citation><mixed-citation xml:lang="en">Sheiko A.M. Language technology toolsin translation quality assurance. Kazan Linguistic Journal. 2023;6(2):282–293. DOI 10.26907/2658-3321.2023.6.2.282-293. (In Russ.)</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
