# Data analytics by complex networks and text mining

**URL:** https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089
**Category:** Feature
**Created:** [March 29, 2017, 7:59am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089 "2017-03-29T07:59:19Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![Renato\_Fabbri](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/renato_fabbri/32/69965_2.png) [@Renato\_Fabbri](https://meta.discourse.org/u/Renato_Fabbri)
#### Post date: [March 29, 2017, 7:59am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/1 "2017-03-29T07:59:19Z")

</div>

Dear users and developers,

by looking through Discourse site, it seems to me that an interface for further data analytics would enhance  
our experience.  
I can tackle this by using complex networks and text mining techniques,  
because I have developed research on these topics.  
Immediate ideas are:  
\*) Derive network by users interactions and vocabulary usage and deliver some graphical interfaces for exploring these structures.  
\*) Counting of words, tags, terms and user activity.

I understand it might be late for coping with GSoC, but if you find it suitable,  
I might apply.  
My apologies for not making this contact earlier, but I handed my doctorate  
dissertation a few days ago and could not concentrate as needed until now.  
Some info about my research and software development efforts are gathered here:

> <https://pastebin.com/iNNuN4fy>

Anyway, this topic might be of use for the Discourse community as a whole and for  
developments outside GSoC.

Best Regards!  
Renato Fabbri

---

<div class="post-metadata">

### Author: ![tophee](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tophee/32/73406_2.png) [@tophee](https://meta.discourse.org/u/tophee)
#### Post date: [March 29, 2017, 8:36am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/2 "2017-03-29T08:36:15Z")

</div>

Personally, I find your thesis interesting as I also intend to do some analysis on my forum at some point in the future. But I am not sure what exactly the feature for discourse would be. Is this supposed to be a tool for admins to be able to identify types of users? If so, what for and don’t you think existing stats are sufficient? Or is it supposed to be a feature for users so they can compare themselves to others like on fitness tracking sites? If so, I again wonder: what for? Aren’t badges and likes enough to allow comparison and perhaps provide some incentive? Maybe I am completely misunderstanding you, so please explain.

---

<div class="post-metadata">

### Author: ![Renato\_Fabbri](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/renato_fabbri/32/69965_2.png) [@Renato\_Fabbri](https://meta.discourse.org/u/Renato_Fabbri)
#### Post date: [March 29, 2017, 9:33am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/3 "2017-03-29T09:33:00Z")

</div>

I thought about the tool for general users  
and think that the usefulness can be glimpsed by  
the following questions:  
\*) Does the current stats make clear what are the all time and recent most active users?  
\*) What users relate to each other beyond what is grasped by browsing through individual topics?  
\*) How does the overall interaction network(s) looks like? What characteristics can a participant take advantage of and how can Discourse encourage fruitful interactions?  
\*) What are the most used words and terms in Discourse and how they relate to each other?  
\*) Do we have interactive and interesting graphical interfaces to the analytics?  
\*) How do linguistic traces differ in user groups and how can it be used by the participants?

Anyway, I think that these kind of analyzes can help users get interested in the legacy  
and admins to showcase or make reports.

---

<div class="post-metadata">

### Author: ![erlend\_sh](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/erlend_sh/32/119475_2.png) [@erlend\_sh](https://meta.discourse.org/u/erlend_sh)
#### Post date: [March 29, 2017, 9:31pm UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/4 "2017-03-29T21:31:06Z")

</div>

I think some really cool data could come out of this, but it’s too experimental as a GSoC project, because it’s hard to tell exactly what we’d get out of it. At the very least we’d need a proof of concept to peak our interest 😉

> [@Renato\_Fabbri](#):
>
> I can tackle this by using complex networks and text mining techniques,

What do you think is _the one_ most interesting piece of data that could be derived from these techniques?

---

<div class="post-metadata">

### Author: ![Renato\_Fabbri](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/renato_fabbri/32/69965_2.png) [@Renato\_Fabbri](https://meta.discourse.org/u/Renato_Fabbri)
#### Post date: [March 30, 2017, 4:12am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/5 "2017-03-30T04:12:23Z")

</div>

The most interesting piece for Discourse IMHO is the interaction network  
because it is simple, informative and eye catching.  
There are a number of proofs of concept(s) in the documentation linked  
in my starting message.  
I will be happy to make some images and measurements from Discourse data,  
if you can send me the database dump or direct me to an interface.  
I might not make a JavaScript interface with D3.js interactive graphs now,  
which are cool to contemplate and useful for investigation.

---

<div class="post-metadata">

### Author: ![Renato\_Fabbri](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/renato_fabbri/32/69965_2.png) [@Renato\_Fabbri](https://meta.discourse.org/u/Renato_Fabbri)
#### Post date: [March 31, 2017, 9:16am UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/6 "2017-03-31T09:16:45Z")

</div>

Just dropping by to see if there really is no feedback on reaching Discourse data for making the proof of concept. Anyway, thank you for your time and nice interaction.

---

<div class="post-metadata">

### Author: ![Mittineague](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mittineague/32/114259_2.png) [@Mittineague](https://meta.discourse.org/u/Mittineague)
#### Post date: [March 31, 2017, 1:35pm UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/7 "2017-03-31T13:35:28Z")

</div>

I have seen topics here over the years where members expressed interest in having various data other than what can be seen on the /users, /dashboard and /report pages. Search here for “statistics” and you can find some.  
AFAIK the typical approach has been to use crafted queries with the [Data Explorer Plugin](https://meta.discourse.org/t/data-explorer-plugin/32566)

There is also this plugin I haven’t tried yet:

> [@Admin Statistics Report](https://meta.discourse.org/t/admin-statistics-report/50943):
>
> Based on the [Admin Statistics Digest spec](https://meta.discourse.org/t/plugin-admin-statistics-digest/44270). Source on GitHub: [GitHub - saiqulhaq/discourse-admin-statistics-digest · GitHub](https://github.com/discourse/discourse-admin-statistics-digest) This plugin sends statistic report like: These users joined some time in the past 60-30 days These are your top non-staff users of the month These formerly very active users are not so active any more The 5 most active responders in the #Support category were: … Most popular topics & posts this month Preview: Email report sample: plugin dashboard: I…

Perhaps the way forward would be to collaborate with @saiqulhaq ?

---

<div class="post-metadata">

### Author: ![saiqulhaq](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saiqulhaq/32/119871_2.png) [@saiqulhaq](https://meta.discourse.org/u/saiqulhaq)
#### Post date: [March 31, 2017, 3:05pm UTC](https://meta.discourse.org/t/data-analytics-by-complex-networks-and-text-mining/60089/8 "2017-03-31T15:05:03Z")

</div>

[Admin Statistic Digest](https://meta.discourse.org/t/admin-statistics-report/50943) plugin generates statistics report in simple way, no multi thread process  
it retrieves data from database, then calculating it, if data retrieval process is more than specified time (maybe 1 or 2 minutes), then the process is aborted

However it would be cool if this feature could be implemented

I am wondering, does text mining process should be executed in the single machine with Discourse server? is it requires high hardware requirement?
