I just set up the following at OpenAI
As Tibo posts on X, Discourse now will automatically update a post that keeps track of the tweets.
How this works?
A Data table as a cache
Discourse Workflows ships with a Data table feature.
This allows us to store structured information between workflow runs.
A critical piece when integrating with X is avoiding API cost overrun. To do that we store a cache of all the information we scraped to avoid re-querying old tweets.
A data table is perfect for this.
A configuration node
“Set fields” is perfect for configuration nodes. In this case I needed some “per workflow” configuration which is used throughout the workflow, topic id to post in, campaign start time, usernames and so on.
It helps keep the workflow easier to reason about.
Http requests to X
The workflow leans on 2 X endpoints:
A very simple one to look up Tibo’s X id.
A second one to look up timeline:
const data = $json;
const checkpoint = $("Read checkpoint").first().json;
const userId = data.user_id || data.data?.id;
if (!userId || !/^[A-Za-z0-9-]+$/.test(userId)) {
throw new Error("Unexpected X user lookup response");
}
const configuration = $("Configuration").first().json;
const apiBase = configuration.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const token = data.next_token || "";
const base = apiBase + "/2/users/" + userId + "/tweets";
const params = "max_results=" + (demo ? "5" : "100") + "&post.fields=id,text,created_at,entities,attachments,note_post&expansions=referenced_posts";
const boundary = checkpoint.cursor ? "since_id=" + encodeURIComponent(checkpoint.cursor) : "start_time=" + encodeURIComponent(configuration.campaign_start_time);
const url = base + "?" + params + "&" + boundary + (token ? "&pagination_token=" + encodeURIComponent(token) : "");
return { user_id: userId, url: url, tweets: data.tweets || [], candidate_cursor: data.candidate_cursor || checkpoint.cursor,
page_count: data.page_count || 0, tokens: data.tokens || [] };
This also starts exposing some implementation details
I built the workflow using an agent and ran a “fake X API” during the process to save on costs (hence the localhost there)
Using flow “If blocks”, we can keep iterating through a users timeline, while we have more pages. A nice implementation detail of X API is that you only pay for actual posts you get, so zero result queries do not cost anything.
Parsing tweets is done using this little script node:
const previous = $("Build page URL").item.json;
const response = $json;
if (!response || !response.meta || !Array.isArray(response.data || [])) {
throw new Error("Incomplete X timeline response; refuse to update the post");
}
if (response.errors?.length) { throw new Error("X returned partial data/errors; refuse to update the post"); }
const next = response.meta.next_token || "";
if (next && (previous.tokens.includes(next) || previous.page_count >= 39)) {
throw new Error("X pagination repeated or exceeded the 40-page safety limit");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
let candidate = previous.candidate_cursor;
for (const tweet of response.data || []) {
if (typeof tweet.id !== "string" || !/^\d+$/.test(tweet.id)) { throw new Error("X post ID must be a decimal string"); }
if (!candidate || greater(tweet.id, candidate)) { candidate = tweet.id; }
}
const normalize = post => {
const note = post.note_post || post.note_tweet;
return { ...post, text: note?.text || post.text, entities: note?.entities || post.entities,
referenced_tweets: post.referenced_posts || post.referenced_tweets || [] };
};
const expanded = Object.fromEntries((response.includes?.posts || response.includes?.tweets || [])
.map(post => [post.id, normalize(post)]));
const page = (response.data || []).map(normalize).map(post => ({ ...post,
referenced_detail_urls: post.referenced_tweets.flatMap(ref => {
const quote = expanded[ref.id];
return quote ? [quote.text || "", ...(quote.entities?.urls || []).map(url => url.unwound_url || url.expanded_url || url.url)] : [];
}) }));
const tweets = previous.tweets.concat(page);
return { user_id: previous.user_id, tweets: tweets, candidate_cursor: candidate, next_token: next,
has_more: !!next, has_tweets: tweets.length > 0, page_count: previous.page_count + 1,
tokens: next ? previous.tokens.concat(next) : previous.tokens };
Adding intelligence
The key ingredient to successfully synchronizing content is intelligence:
We pass in a list of tweets as input → we get a properly extracted JSON document out. This kind of process requires intelligence.
It both formats the content AND finds the most relevant pieces of information.
Updating the Discourse post
We use a script node to render the table per:
const extracted = $("Poll result").first().json;
const checkpoint = $("Read checkpoint").first().json;
const apiBase = $("Configuration").first().json.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const raw = String($json.post?.raw || "");
const start = "<!-- x-ships-table:start -->";
const end = "<!-- x-ships-table:end -->";
const hasStart = raw.includes(start);
const hasEnd = raw.includes(end);
if (hasStart !== hasEnd ||
(hasStart && (raw.indexOf(end) < raw.indexOf(start) ||
raw.indexOf(start, raw.indexOf(start) + start.length) !== -1 ||
raw.indexOf(end, raw.indexOf(end) + end.length) !== -1))) {
throw new Error("Ambiguous workflow-owned table markers in designated post");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
if (checkpoint.cursor && extracted.candidate_cursor && greater(checkpoint.cursor, extracted.candidate_cursor)) {
throw new Error("Refusing to move the timeline cursor backwards");
}
const bySlot = {};
for (const ship of checkpoint.ships.concat(extracted.ships)) {
if (!Number.isInteger(ship.day) || ship.day < 1 || ship.day > 28 ||
!new RegExp("^" + ship.day + "(?:\\.[1-9][0-9]?)?$").test(ship.slot) ||
typeof ship.source_id !== "string" || !/^\d+$/.test(ship.source_id) ||
typeof ship.announcement !== "string" || !ship.announcement || ship.announcement.length > 400 ||
(ship.details_url && !/^https?:\/\/[^\s<>()[\]|\\]+$/.test(ship.details_url))) {
throw new Error("Invalid stored ship slot");
}
if (!bySlot[ship.slot] || greater(ship.source_id, bySlot[ship.slot].source_id)) { bySlot[ship.slot] = ship; }
}
const escapeCell = value => String(value).replace(/[|\r\n]/g, " ").replace(/\\/g, "\\\\").replace(/([\[\]_*`])/g, "\\$1");
const rows = [];
for (let day = 1; day <= 28; day++) {
const ships = Object.values(bySlot).filter(ship => ship.day === day)
.sort((a, b) => a.slot.localeCompare(b.slot, undefined, { numeric: true }));
if (!ships.length) { rows.push("| " + day + " | Pending | — | — | — |"); }
else for (const ship of ships) {
const provenance = demo
? "[Demo post](" + apiBase + "/demo/posts/" + encodeURIComponent(ship.source_id) + ") (synthetic)"
: "[X post](https://x.com/" + encodeURIComponent($("Configuration").first().json.username) + "/status/" + encodeURIComponent(ship.source_id) + ")";
const details = ship.details_url ? "[Details](" + ship.details_url + ")" : "—";
rows.push("| " + day + " | " + escapeCell(ship.slot) + " | " + escapeCell(ship.announcement) + " | " + details + " | " + provenance + " |");
}
}
const ships = Object.values(bySlot).sort((a, b) => a.day - b.day || a.slot.localeCompare(b.slot, undefined, { numeric: true }));
const updated = start + "\n| Day | Ship | Announcement | Details | Source |\n| --- | --- | --- | --- | --- |\n" + rows.join("\n") + "\n" + end;
const replacement = hasStart
? raw.slice(0, raw.indexOf(start)) + updated + raw.slice(raw.indexOf(end) + end.length)
: raw + (raw ? (raw.endsWith("\n") ? "\n" : "\n\n") : "") + updated;
return { changed: replacement !== raw, raw: replacement, ships_json: JSON.stringify(ships),
since_id: extracted.candidate_cursor, user_id: extracted.user_id };
The small trick is that the Discourse post has a special marker:
That way after we generate the table we can double if anything changed prior to updating the table.
If the table changes we update the tracking Data Table:
How I built this?
Workflows are incredibly powerful features building giant ones like this by hand would take a long time.
Instead of doing that I leaned on an agent, I use term-llm.com for all my development, it integrates very cleanly into dv. Of course your agent of choice, but a key feature here is having a working Discourse environment.
My steps were:
-
dv new workflow-build- to create a new clean environment.- I use a special dv hook that fires up a term-llm serve web service that connects back to my hub, this gives me an instant “rich agent” that is in a throwaway environment I can run wild with.
-
I prompted it:
I used GPT 6.1 Sol for the build and already had something going after a single shot:
- I refined it
-
The first attempt was not using a data table - I prompted to move storage there
-
Next I iterated to move all configuration into a central node so it would be easier to manage it
-
I exported the workflow - and then imported into production
-
I iterated a bit with the agent cause in production there were a few missing edge cases (our export xml does not include data table definitions and agent definitions)
What if you need to build something like this?
This post will be incredibly helpful to agents, a simple prompt along the lines of:
Synchronize important discussion about “X” on Accounts A,B,C into my topic following similar method in LINK TO META
Will allow you to do arbitrary data synchronization from X.
Social → Discourse is a very important vector for community, workflows allow you to keep a long running record that is open to more nuanced discussion without forcing existing people in your company to centralize all posting of content on Discourse.
Workflows are the building block you can use for this kind of synchronization and the options here are quite wide.
- Post shadow topics when specific people post content on social
- Keep a long running topic up to date based on activity on social
- … and much more
















