UrbanPro
true

Learn Data Science from the Best Tutors

  • Affordable fees
  • 1-1 or Group class
  • Flexible Timings
  • Verified Tutors

Search in

Topic Modeling in Text Mining : LDA

Ashish R.
13/05/2017 0 0

Latent Dirichlet allocation (LDA)

Topic modeling is a method for unsupervised classification of text documents, similar to clustering on numeric data, which finds natural groups of items even when we’re not sure what we’re looking for. In clustering one entity can belongs to one group only, whereas in topic modeling a word can belongs to multiple groups/clusters with varying level of probability. The input of the model is a text document/ or a set of documents. The out of the model is to split the documents into multiple K groups and then determining a topic from each group based on the association of the most important words in the respective group. The number of topic which is equivalent to the number of clusters in cluster analysis (K) has to be selected based various heuristics on how many topics might be extracted from the document/s. LDA treats each document as a mixture of topics, and each topic as a mixture of words. This allows documents to “overlap” with each other in terms of content, rather than being separated into discrete groups.

 As an output of LDA model, if we decide to find out K topics then our set of documents are segregated into K groups. The key words or the tokens in each group receive a beta value describing how strong the tokenized word is associated with many other words (tokens) within the group. The larger the value of beta explains the more importance of the word in that group. Top 6-10 words with the largest beta values are chosen to decide the topic that is depicting by that group of words. The topic is decided based on human intelligence on understanding the meaning of those words in the underlying context of the collected documents.

How to determine the number of topic from a set of documents

Hierarchical clustering analysis is performed on the group of words that are collected from the corpus to determine the number of clusters to form. Using distance metric like Levenshtein distance, Hamming Distance etc., the distance among the words are plotted in a dendrogram. The vertical axis of the dendrogram scales the chosen distance metric. Based on the word cloud formation, we decide what distance to consider as a cut off distance to determine the number of appropriate groups to be formed with the set of documents. This is similar like hierarchical clustering with numeric data values where usually Euclidean distance is considered by default. 

 

0 Dislike
Follow 0

Please Enter a comment

Submit

Other Lessons for You

TOP 10 Tools for Data Science
TOP 10 Tools for Data Science1. Python2. SQL3. R4. Tableau5. PowerBI6. Java7. Julia8. Scala9. SAS10. ExcelTOP 10 Websites for Data Science1. Coursera3. EdX4. Udacity5. Kaggle6. Analytics Vidhya7. KDNuggets8....

Data Science: Case Studies
Modules Training Practice Case Studies Module 2: Data Visualization and Summarization 10 15 1. Crime Data 2. Depression & anxiety 3....

Regularisation in Machine Learning
Regularization In Machine Learning, Regularization is the concept of shrinking or regularizing the coefficients towards zero. It helps the model to prevent overfitting. Overfitting in Machine Learning...

R vs Statistics
I frequently asked the below question from my students: 'Do I You need Statistics to learn R Programming?' The answer is, NO. If you want to learn R programming only, Stat is not required. You can be...

Learn Data Science In 8 Steps
8 Steps To Learn Data Science There have been a lot of surveys over the past few years on the educational background of data scientists. As a result, there have also been many different results. In the...
X

Looking for Data Science Classes?

The best tutors for Data Science Classes are on UrbanPro

  • Select the best Tutor
  • Book & Attend a Free Demo
  • Pay and start Learning

Learn Data Science with the Best Tutors

The best Tutors for Data Science Classes are on UrbanPro

This website uses cookies

We use cookies to improve user experience. Choose what cookies you allow us to use. You can read more about our Cookie Policy in our Privacy Policy

Accept All
Decline All

UrbanPro.com is India's largest network of most trusted tutors and institutes. Over 55 lakh students rely on UrbanPro.com, to fulfill their learning requirements across 1,000+ categories. Using UrbanPro.com, parents, and students can compare multiple Tutors and Institutes and choose the one that best suits their requirements. More than 7.5 lakh verified Tutors and Institutes are helping millions of students every day and growing their tutoring business on UrbanPro.com. Whether you are looking for a tutor to learn mathematics, a German language trainer to brush up your German language skills or an institute to upgrade your IT skills, we have got the best selection of Tutors and Training Institutes for you. Read more