# \[bounty\] Google+ (private ) communities: export screenscraper + importer

**URL:** https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029
**Category:** Marketplace
**Created:** [January 31, 2019, 3:40pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029 "2019-01-31T15:40:04Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![adorfer](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/adorfer/32/82089_2.png) [@adorfer](https://meta.discourse.org/u/adorfer)
#### Post date: [January 31, 2019, 3:40pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/1 "2019-01-31T15:40:04Z")

</div>

Continuing the discussion from [Importing from Google Groups](https://meta.discourse.org/t/importing-from-google-groups/76860):

I am looking for a way to salvage a existing GooglePlus Communities.

- groups ares “closed” (non public) ones: Not (fully) exportable the tools i know  
in other words: export has to run in the user context.  
(most drastic approach would be _“browser-screenscraping via Silenium”_)
- Import into discourse should include images
- include comments on the posts and comments on individual images
- preserving as many detail as possible

**Bounty:**

- I would put via paypal 200€ or 200USD (as you like ) on the table.
- this will probably not cover the cost of a full development
- idea is that more people join
- code will be made public available. (hopefully that does not turn people away from joining)

Export-screenscraper has to meet the existing timeline:

> **[Google Workspace Updates: The future of Currents and the next generation of...](https://workspaceupdates.googleblog.com/2022/02/currents-spaces-migration.html?visit_id=639159096077870738-3301967640&rd=1)**

(in other words: be operational latest within 1 months / end of february 2019)

---

<div class="post-metadata">

### Author: ![erlend\_sh](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/erlend_sh/32/119475_2.png) [@erlend\_sh](https://meta.discourse.org/u/erlend_sh)
#### Post date: [January 31, 2019, 4:11pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/3 "2019-01-31T16:11:32Z")

</div>

We’d be happy to put up another €200 for this, payable only via PayPal.

---

<div class="post-metadata">

### Author: ![notriddle](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/notriddle/32/133055_2.png) [@notriddle](https://meta.discourse.org/u/notriddle)
#### Post date: [January 31, 2019, 10:08pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/4 "2019-01-31T22:08:56Z")

</div>

I’ll add another $200, again paid via PayPal.

---

<div class="post-metadata">

### Author: ![Stephane\_Buisson](https://avatars.discourse-cdn.com/v4/letter/s/87869e/32.png) [@Stephane\_Buisson](https://meta.discourse.org/u/Stephane_Buisson)
#### Post date: [February 2, 2019, 11:20am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/5 "2019-02-02T11:20:52Z")

</div>

Hi Guys,

Some tools are available see below, but From G+ to Owner/admin we should wait early March.  
"Google+ Communities  
To download data for Communities where you’re an owner or moderator, select Google+ Communities. You will get:

Names and links to Google+ profiles of community owners, moderators, members, applicants, banned members, and invitees  
Links to posts shared with the community  
Community metadata, including community picture, community settings, content control settings, your role, and community categories  
Important: Starting early March 2019, you will also be able to download additional details from public communities, including author, body, and photos for every community post."

Now here the tool I am thinking about : [https://gplus-exporter.friendsplus.me/](https://gplus-exporter.friendsplus.me/)

Stéphane is owner of [Google Workspace Updates: New community features for Google Chat and an update on Currents](https://plus.google.com/communities/118113483589382049502) and also interested in Discourse.

---

<div class="post-metadata">

### Author: ![pfaffman](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/pfaffman/32/120154_2.png) [@pfaffman](https://meta.discourse.org/u/pfaffman)
#### Post date: [February 2, 2019, 9:04pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/6 "2019-02-02T21:04:20Z")

</div>

But it looks like you don’t get email addresses, so it’ll be difficult to allow users to be connected with their posts.

I’ll have a look at the tool linked above. $600 is still short of what I generally charge for a new importer (a recent project was over $3000) , and it looks like it’ll have a short window of usefulness.

I can have a look next week at the resources linked and got much time they’ll save.

---

<div class="post-metadata">

### Author: ![adorfer](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/adorfer/32/82089_2.png) [@adorfer](https://meta.discourse.org/u/adorfer)
#### Post date: [February 3, 2019, 6:57pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/7 "2019-02-03T18:57:21Z")

</div>

> [@pfaffman](#):
>
> But it looks like you don’t get email addresses, so it’ll be difficult to allow users to be connected with their posts.

This is bit of a hazzle. An admin needs to join/merge the accounts “to be claimed” later manually.  
But at least for Groups with just a few dozend really active members it’s feasable.

> [@Stephane\_Buisson](#):
>
> Now here the tool I am thinking about : [https://gplus-exporter.friendsplus.me/](https://gplus-exporter.friendsplus.me/)

the google exporter may be good starting point and even reduce a lot of preassure “to get it completed fast”.  
(But i have not looked into the format of the G+exporter yet, i am not shure if all neccesary/relevant details is covered. Now would be the time perhaps to get in contact with the author in order to ask for supplemental data to fetch, especially since there seem to be nearly daily releases.)

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 19, 2019, 12:14am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/8 "2019-02-19T00:14:58Z")

</div>

Just FYI, I’ve been hacking at this for the K40 community and maybe some other communities, starting from a [friendsplus.me](http://friendsplus.me) JSON export, and didn’t see this post until a moment ago. I have used the [friendsplus.me](http://friendsplus.me) exporter to create static Jekyll archives of several communities, so I’m already up to speed with their JSON format.

I’ve never touched ruby before, so I’m learning Discourse and ruby at the same time, but it doesn’t look too bad. I took the route of creating suspended `@example.com` (so that they cannot be validated by an attacker) fake users for admins to merge later, though I’ll probably add the ability to provide a JSON map to known existing users at the time of the import.

It looks like the ning importer has dealt with many of the same general needs, so I’ve been reading it as a pattern for this work.

I’m _not_ seeking the bounty and I _do_ intend to share my work. I intend to be involved with running it a limited number of times and intend not to be a long-term maintainer for the script. I intend to offer a PR when it works, but given the limited lifetime of the script it might be reasonable not to merge the PR, especially if it’s not up to quality expectations, and instead leave it as documentation.

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 19, 2019, 12:24am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/9 "2019-02-19T00:24:35Z")

</div>

Also, there are at least two other possible source of data for G+ info to be imported into Discourse.

1. Google has promised (¯\_(ツ)_/¯) to provide a community takeout option scant days before starting to delete the entire site. No schema has been documented or discussed for this data.
2. There is an [open source migration tool](https://github.com/RomainVialard/Google-Plus-Community-Migrator) that could probably be “borrowed” — the K40 community tried this tool and it did not succeed, but they have over 5000 posts. It might work for smaller communities, either as a long term move directly, or as a way to preserve a community while working on an import from their appscript clone into Discourse for a more long-term solution.
3. [Update: added] [GitHub - FiXato/Plexodus-Tools: A collection of tools to process the Google Plus-related data from Google Takeout. · GitHub](https://github.com/FiXato/Plexodus-Tools) — also open source — has code both for working with takeout and talking to the G+ API (which has about two weeks of intermittent life left)

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 20, 2019, 2:41am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/10 "2019-02-20T02:41:57Z")

</div>

I should be clearer: If someone else wants to pursue the bounty, such as it is, I’m happy for you to do it and will contribute what I’ve learned so far towards the project, without asking to share the bounty. I just want this to exist.

> [@mcdanlj](#):
>
> I took the route of creating suspended `@example.com` (so that they cannot be validated by an attacker) fake users for admins to merge later,

I see that `GoogleUserInfo` has `google_user_id` on which I suppose I can join for G+ users who have already logged in via google auth. That’s the same ID that’s in the export.

I’m wondering if I could set just the `user_id` and `google_user_id` fields in `GoogleUserInfo` for the fake users I create because they haven’t logged in yet via google auth, and then if they log in later via google auth, the `google_oauth2_authenticator` will merge the users automatically and fix their user name and email from the oauth2 response?

---

<div class="post-metadata">

### Author: ![pfaffman](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/pfaffman/32/120154_2.png) [@pfaffman](https://meta.discourse.org/u/pfaffman)
#### Post date: [February 20, 2019, 4:00am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/11 "2019-02-20T04:00:19Z")

</div>

If you can share enough data that I can write an importer, I can see what it would take.

I would be seeking the bounty. This is my day job. I would also offer to run imports for people with budgets, as each one is surprisingly different. Once I’ve run an import to get the code to work with one site, I’d share the code. I’ve got at least one person interested in an import, which could make it worth my while.

I’ve written several importers.

I will share the code.

There is no issue with long term maintenance! This game will be over in a few weeks.

If you can get me the dump you have I can see about helping match those Google ids. I think that what you suggest might be possible, but I’d need to see the data to tell for sure.

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 20, 2019, 11:38am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/12 "2019-02-20T11:38:01Z")

</div>

This morning, I read ActiveRecord docs, so `GoogleUserInfo.find_by` looks like my friend. I’m a Go/Python/C developer so I just have to look everything up as I go along… ☺

I have a local test environment running in which to do test imports, so that’s not so hard.

Here’s the information I have outside Discourse so far:

1. The author of the Friends+Me exporter has a commented exemplar for schema doc at [https://docs.google.com/document/d/1gOYJe61sI1GbO9qpFZJtwY3vdsZalsxtPgUnB2cvzAw/edit](https://docs.google.com/document/d/1gOYJe61sI1GbO9qpFZJtwY3vdsZalsxtPgUnB2cvzAw/edit)
2. [Anthony Bolgar / K40 · GitLab](https://gitlab.com/funinthefalls/k40) is an import I did into a Jekyll site before I realized that this was really an option, with the [the exported feed JSON](https://gitlab.com/funinthefalls/k40/raw/master/import/feed.json), [a mapping of URLs to image files in the repository](https://gitlab.com/funinthefalls/k40/raw/master/import/map.json?inline=false), [the python script I hacked together to build the jekyll site](https://gitlab.com/funinthefalls/k40/raw/master/bin/import-friendsplus?inline=false), and [4807 images checked into the repository](https://gitlab.com/funinthefalls/k40/tree/master/images).

I’m having fun hacking at this for a bit. If you get to the point where you are ready to actually start work on it, drop me a PM and I’ll provide what I have so far, though I don’t promise to stop playing with it myself at that point. I don’t want to use this topic as an ugly form of source code management. 😉

I expect that outside of markup, most of the code will be re-usable for importing google community takeout archives when google actually releases them, so this might be a head start on being able to do more imports for people who have just waited until the last minute.

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 20, 2019, 11:51am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/13 "2019-02-20T11:51:34Z")

</div>

> [@pfaffman](#):
>
> I think that what you suggest might be possible

Here’s my untested attempt to import users:

```plaintext
    def import_author_user(author)
      id = author["id"]
      if not @users[id].present?
        google_user_info = ::GoogleUserInfo.find_by(google_user_id: author["id"]
        if google_user_info.nil?
          name = author["name"]
          email = "gplus-#{name.gsub(/\s+/, "")}-#{id}@example.com"
          {
            id: id,
            email: email,
            name: name,
            post_create_action: proc do |newuser|
              newuser.approved = true
              newuser.approved_by_id = @system_user.id
              newuser.approved_at = newuser.created_at
              newuser.save
              ::GoogleUserInfo.create({ 
                user_id: newuser.id,
                google_user_id: id,
              })
            end
          }
        else
          email = google_user_info.email
        end
        @users[id] = email
      end
    end
  end

```

I explicitly intend by `@example.com` to prevent hostile takeover, and for the `google_user_id` to allow later automatic user merge when they log in.

Again, I have no idea if that will work.

---

<div class="post-metadata">

### Author: ![gerhard](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/gerhard/32/119479_2.png) [@gerhard](https://meta.discourse.org/u/gerhard)
#### Post date: [February 20, 2019, 12:05pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/14 "2019-02-20T12:05:29Z")

</div>

> [@mcdanlj](#):
>
> I explicitly intend by `@example.com` to prevent hostile takeover

I suggest you use a `.invalid` domain in order to prevent outgoing emails. Something like this:

```ruby
email = "#{id}@gplus.invalid"

```

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 20, 2019, 12:34pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/15 "2019-02-20T12:34:03Z")

</div>

> [@gerhard](#):
>
> I suggest you use a `.invalid` domain

That’s a much better idea, I had forgotten about `.invalid` and it’s exactly right for the purpose. Thank you!

(`example.com` shouldn’t result in outgoing emails either by specification, but its purpose is primarily documentation.)

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 21, 2019, 3:49am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/16 "2019-02-21T03:49:52Z")

</div>

I think I’ve worked out how to do message formatting. I expect that the `Post.cook_methods[:regular]` will be able to handle embedded bits of html like `<b>` `<i>` and `<s>` which I found I really needed to use with my jekyll import because nesting doesn’t work the same between markdown and what G+ produced.

The only things I know left for message formatting are:

1. Turning plus-references into at-references that will resolve internally; there can be references outside the community to users not imported, so I’ll have to only conditionally turn them into at-references if the user exists on the system already or in the import. Update: I think I have this figured out. There might be a better way than `"<a class="mention" href="/u/#{user.name}">@#{user.name}</a>"` though?
2. Coming up with a reasonable title for a topic from a post. I have an algorithm but I’ll have to validate that it generates OK titles.

---

<div class="post-metadata">

### Author: ![gerhard](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/gerhard/32/119479_2.png) [@gerhard](https://meta.discourse.org/u/gerhard)
#### Post date: [February 21, 2019, 12:01pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/17 "2019-02-21T12:01:24Z")

</div>

> [@mcdanlj](#):
>
> I think I have this figured out. There might be a better way than `"<a class="mention" href="/u/#{user.name}">@#{user.name}</a>"` though?

Using `@username` should be enough. Discourse will create the proper HTML when the post gets cooked. I guess you figured out a way to find existing users by their plus-reference?

BTW: Are you writing a proper import script by using our [base importer](https://github.com/discourse/discourse/blob/master/script/import_scripts/base.rb)? Depending on how you import the users, you should be able to use `find_user_by_import_id(google_plus_user_id)` to find existing users.

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 21, 2019, 12:40pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/18 "2019-02-21T12:40:16Z")

</div>

> [@gerhard](#):
>
> Discourse will create the proper HTML when the post gets cooked.

Oh, great! That’s easier. Glad to know I was trying too hard.

I do want to be able to “fix up” references later to users who don’t exist, so if I don’t find them either in the import or on the system, I was thinking of something like `"<b data-google-plus-id="123456789">+GooglePlusName</b>"`

I’m using the base importer, using ning.rb as a pattern since it seemed to be the closest analog. I’m using `find_user_by_import_id` but not all users will be in the import; I’ll have multiple imports across different G+ communities and many people will have already signed in so I also need to look in `GoogleUserInfo` to resolve the mapping. My plan is actually to import all the users across all the imports, and then to import the posts across all the imports, so that cross-import user references are all resolved correctly. A lot of the users interacted in a lot of related communities on G+ and I want to create a familiar experience for them.

Now, whether it’s a “proper” import script others can judge! 😉 Thanks for putting up with half-formed thoughts from this new-to-Discourse-and-new-to-Ruby guy, and for help pointing me in the right direction. I’ll say that it’s clear to me that some real attention has been paid to making imports work in Discourse, which is motivational for trying to add an importer!

Thanks!

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 22, 2019, 12:16pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/19 "2019-02-22T12:16:24Z")

</div>

@gerhard — A related question… I now notice that the ning.rb importer has special handling for youtube links. Should I expect that youtube (and vimeo, etc?) links will be recognized and transformed into iframes with a viewer when the post is cooked, just like at-references to users? (The python code I wrote for static site imports has substantial regexp handling for youtube that I could pretty much lift intact, but if Discourse does it better I’d rather just trust Discourse to do the right thing.)

---

<div class="post-metadata">

### Author: ![gerhard](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/gerhard/32/119479_2.png) [@gerhard](https://meta.discourse.org/u/gerhard)
#### Post date: [February 22, 2019, 12:28pm UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/20 "2019-02-22T12:28:49Z")

</div>

The Ning importer removes iframes so that Discourse can detect the links and create oneboxes. You don’t have to do anything special in your import script as long as the link is on a line itself.

[https://meta.discourse.org/t/what-is-a-onebox/78060](https://meta.discourse.org/t/what-is-a-onebox/78060)

---

<div class="post-metadata">

### Author: ![mcdanlj](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/mcdanlj/32/131829_2.png) [@mcdanlj](https://meta.discourse.org/u/mcdanlj)
#### Post date: [February 24, 2019, 2:58am UTC](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029/21 "2019-02-24T02:58:53Z")

</div>

Progress report: I have written ~~something vaguely resembling ruby~~ an importer that looks like it covers all the current requirements in general design. (It doesn’t translate G+ +1’s into likes, because the exporter does not represent the +1s.) I haven’t ~~run it at all~~ actually imported data with it yet. 🙄

My script takes as arguments paths to files containing Friends+Me Google+ Exporter JSON export files, Friends+Me Google+ Exporter CSV image map files, and a single JSON file that maps Google+ Category IDs to Discourse categories, subcategories, and tags. It does not create new categories; if a category is missing it complains and bails out after writing a new file in which to fill the missing information to complete the import. It expects all the categories to have been created already.

The idea is to do all the imports into a single discourse in one pass, so that all possible “plus-mentions” of other G+ users across multiple G+ communities turn into “at-mentions” in Discourse, even for people not active in the community in which they are mentioned, as long as they wrote some post or comment somewhere in the _whole set_ of data being imported. This is because so far it looks like I’ll be importing about 10 communities, with about 300MB of input JSON and about 40GB of images.

It is intended to work on a Discourse instance that already has users referenced in the import by google ID, and that already has content and categories created. I hope that it will also make it possible for people to log in with google OAuth2 after the import and automatically own their content because their google auth ID is tied to the fake account holding the data, so that their ability to own their own content is preserved.

I expect that a 431-line file that has never seen an interpreter will be _loads of fun_ to debug, especially when written and being tested by someone who has never written any ruby before. I don’t pretend that writing this script is the largest part of the work. I’ll share it now or any later time with anyone seeking the bounty, as long as you’ll share your fixes with me regardless of bounty progress; just PM me. I’ll share it myself under GPLv3 at such time as I get it working. In the meantime, I’m considering this work my contribution toward **someone else** claiming the bounty, to make it more likely to be worth the time for whoever takes it to completion, because of the comment above that the bounty is smaller than typical.

[Next page](https://meta.discourse.org/t/bounty-google-private-communities-export-screenscraper-importer/108029.md?page=2)
