loading
Papers Papers/2022 Papers Papers/2022

Research.Publish.Connect.

Paper

Author: Johannes Schneider

Affiliation: University of Liechtenstein, Vaduz, Liechtenstein

Keyword(s): Topic Modeling, Sentence Embeddings, Bag of Sentences.

Abstract: Pre-trained language models have led to a new state-of-the-art in many NLP tasks. However, for topic modeling, statistical generative models such as LDA are still prevalent, which do not easily allow incorporating contextual word vectors. They might yield topics that do not align well with human judgment. In this work, we propose a novel topic modeling and inference algorithm. We suggest a bag of sentences (BoS) approach using sentences as the unit of analysis. We leverage pre-trained sentence embeddings by combining generative process models and clustering. We derive a fast inference algorithm based on expectation maximization, hard assignments, and an annealing process. The evaluation shows that our method yields state-of-the art results with relatively little computational demands. Our method is also more flexible compared to prior works leveraging word embeddings, since it provides the possibility to customize topic-document distributions using priors. Code and data is at https:/ /github.com/JohnTailor/BertSenClu. (More)

CC BY-NC-ND 4.0

Sign In Guest: Register as new SciTePress user now for free.

Sign In SciTePress user: please login.

PDF ImageMy Papers

You are not signed in, therefore limits apply to your IP address 18.117.107.93

In the current month:
Recent papers: 100 available of 100 total
2+ years older papers: 200 available of 200 total

Paper citation in several formats:
Schneider, J. (2024). Efficient and Flexible Topic Modeling Using Pretrained Embeddings and Bag of Sentences. In Proceedings of the 16th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART; ISBN 978-989-758-680-4; ISSN 2184-433X, SciTePress, pages 407-418. DOI: 10.5220/0012404000003636

@conference{icaart24,
author={Johannes Schneider.},
title={Efficient and Flexible Topic Modeling Using Pretrained Embeddings and Bag of Sentences},
booktitle={Proceedings of the 16th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART},
year={2024},
pages={407-418},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0012404000003636},
isbn={978-989-758-680-4},
issn={2184-433X},
}

TY - CONF

JO - Proceedings of the 16th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART
TI - Efficient and Flexible Topic Modeling Using Pretrained Embeddings and Bag of Sentences
SN - 978-989-758-680-4
IS - 2184-433X
AU - Schneider, J.
PY - 2024
SP - 407
EP - 418
DO - 10.5220/0012404000003636
PB - SciTePress