# Improving Discourse static HTML archive

**URL:** https://meta.discourse.org/t/improving-discourse-static-html-archive/112497
**Category:** Feature
**Created:** [March 25, 2019, 12:51pm UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497 "2019-03-25T12:51:32Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![saurabhp](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saurabhp/32/125182_2.png) [@saurabhp](https://meta.discourse.org/u/saurabhp)
#### Post date: [March 25, 2019, 12:51pm UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/1 "2019-03-25T12:51:32Z")

</div>

It is recommended to use HTTrack to take a dump of static HTML and host that as a static archived website. But the layout for crawlers is not very pretty to host it as a static site. I will be working on improving the layout and adding necessary data to the static website. You can see the crawler layout at [https://meta.discourse.org/?_escaped\_fragment_](https://meta.discourse.org/?_escaped_fragment_) which I will try to improve.

This is just a placeholder to link with changes I make so that someone reviewing it can get more context.

Let me know if you have any suggestions on this topic.

Thanks

---

<div class="post-metadata">

### Author: ![saurabhp](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saurabhp/32/125182_2.png) [@saurabhp](https://meta.discourse.org/u/saurabhp)
#### Post date: [March 29, 2019, 12:33pm UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/2 "2019-03-29T12:33:12Z")

</div>

I have created few pull requests related to this and added screenshots in them:  
[https://github.com/discourse/discourse/pull/7250](https://github.com/discourse/discourse/pull/7250)  
[https://github.com/discourse/discourse/pull/7270](https://github.com/discourse/discourse/pull/7270)  
[https://github.com/discourse/discourse/pull/7286](https://github.com/discourse/discourse/pull/7286)

Let me know if you have any suggestions.

---

<div class="post-metadata">

### Author: ![tgxworld](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tgxworld/32/106117_2.png) [@tgxworld](https://meta.discourse.org/u/tgxworld)
#### Post date: [April 2, 2019, 5:02am UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/3 "2019-04-02T05:02:09Z")

</div>

> [@saurabhp](#):
>
> It is recommended to use HTTrack to take a dump of static HTML and host that as a static archived website.

Sorry in advanced for my question since I’m not very familiar with HTTrack. Why do we need to use HTTrack to take a dump of the static HTML page and host that as a static archived website?

---

<div class="post-metadata">

### Author: ![saurabhp](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saurabhp/32/125182_2.png) [@saurabhp](https://meta.discourse.org/u/saurabhp)
#### Post date: [April 2, 2019, 6:29am UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/4 "2019-04-02T06:29:19Z")

</div>

Hey,  
You can go through these links to get more context related to this:

> [@A basic Discourse archival tool](https://meta.discourse.org/t/a-basic-discourse-archival-tool/62614/48):
>
> Hi all – just jumping in here to say that @mcmcclur’s code was exactly what I was looking for! So thank you very much for sharing slight_smile I made a few small modifications (mainly additional code that makes sure to grab all posts in a topic, not just the first twenty) and the code is here: [GitHub - kitsandkats/ArchiveDiscourse: Code for archiving a Discourse site into static HTML.](https://github.com/kitsandkats/ArchiveDiscourse), forked from @mcmcclur’s original repo and stored as a python file instead of a Jupyter notebook. I’m very h…

> [@Archive an old forum "in place" to start a new Discourse forum](https://meta.discourse.org/t/archive-an-old-forum-in-place-to-start-a-new-discourse-forum/13433):
>
> When possible, we recommend archiving old forums rather than importing them into Discourse. Sometimes this isn’t possible, for various reasons, but getting a fresh start and moving into a new community platform without carrying all your old baggage across is often appealing. [image] Maybe your community needs a reboot. But how do you: Keep the valuable information at old forums around without keeping that ancient forum software running, too, with all its future security vulnerabilities? A…

HTTrack will basically just crawl your website and create a static HTML dump which you can host as a static website.

Quoting from the link above on why people want it.

> [@A basic Discourse archival tool](https://meta.discourse.org/t/a-basic-discourse-archival-tool/62614/1):
>
> I’ve been using Discourse for a couple of years now as a discussion board when teaching my college math classes so, every few months, I retire one or two sites and start one or two more. Obviously, the discussions on the retiring sites have value so I really needed some way to save them

Let me know if you have any other questions.

---

<div class="post-metadata">

### Author: ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)
#### Post date: [April 2, 2019, 6:30am UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/5 "2019-04-02T06:30:49Z")

</div>

You do not “need” to use httrack tool you can use recursive wget and other similar command line Linuxy spidering tools as well.

---

<div class="post-metadata">

### Author: ![saurabhp](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/saurabhp/32/125182_2.png) [@saurabhp](https://meta.discourse.org/u/saurabhp)
#### Post date: [April 7, 2019, 6:33am UTC](https://meta.discourse.org/t/improving-discourse-static-html-archive/112497/6 "2019-04-07T06:33:12Z")

</div>

Just an update regarding this.

All 3 pull requests have been merged. I’m adding screenshots with the new static archive look here below. Let me know if any of you have any suggestions on things to improve.

 ![Screenshot%20from%202019-04-07%2012-03-07](https://global.discourse-cdn.com/meta/original/3X/e/b/eb66311e6c9a48d8b6603d25e6ba6370b65c22fc.png) ![Screenshot%20from%202019-04-07%2012-02-29](https://global.discourse-cdn.com/meta/original/3X/e/b/eba97194a890d823eff96528048f5c6add61e01b.png) ![Screenshot%20from%202019-04-07%2012-02-06](https://global.discourse-cdn.com/meta/original/3X/f/1/f171510a81e956d2f509d8da54e75f0c8af43d57.png) ![Screenshot%20from%202019-04-06%2010-54-40](https://global.discourse-cdn.com/meta/original/3X/c/f/cfa14c4e24bd8f081d40fbb1ea46f4d5e6018108.png)
