Data mining is a popular research area that has been studied by many researchers and focuses on finding unforeseen and important information in large dataset. Social media data is one of the most popular and large heterogeneous data collected from social networking sites, microblogs, photo or video sharing sites. Social media represents the entities and their relations. One of the popular data structures used to represent large heterogeneous data in the field of data mining is graphs. The nodes of a graph represent entities and the edges of a graph represent the relations between the entities. So, graph mining is one of the most popular subdivisions of data mining. A frequent pattern is referred to as pattern that is more frequently encountered than the user-defined threshold in a dataset. Frequent patterns in a dataset can give important information about dataset. Using this information, data can be classified or clustered. Frequent patterns can provide different perspective on social media data with respect to sociology, consumer behaviour, marketing, communities. In this thesis, popular frequent pattern mining algorithms have been examined and it has been observed that most algorithms are not suitable for large datasets. Since data in today’s world, especially social networks, has very large data, the existing pattern mining algorithms are not suitable for this data. The aim of this thesis is to implement an existing frequent pattern mining algorithm in parallel manner and to find frequent patterns in a social media data.
