根据 X 动态自动更新 Discourse 帖子

我刚刚在 OpenAI 设置了以下内容

正如 Tibo 在 X 上发布的,Discourse 现在会自动更新一个帖子,该帖子用于跟踪推文。

这是如何运作的?

作为缓存的数据表

Discourse Workflows 自带一个数据表(Data table)功能。

这允许我们在工作流运行之间存储结构化信息。

与 X 集成时的一个关键部分是避免 API 成本超支。为此,我们存储了所有抓取信息的缓存,以避免重复查询旧的推文。

数据表非常适合此用途。

配置节点

“设置字段”(Set fields)非常适合用作配置节点。在这种情况下,我需要一些“每个工作流”的配置,这些配置在整个工作流中使用,例如要发布帖子的主题 ID、活动开始时间、用户名等。

这有助于使工作流更容易理解。

向 X 发送 HTTP 请求

该工作流依赖于 2 个 X 端点:

一个非常简单的端点,用于查找 Tibo 的 X ID。

第二个用于查找时间线:

const data = $json;
const checkpoint = $("Read checkpoint").first().json;
const userId = data.user_id || data.data?.id;
if (!userId || !/^[A-Za-z0-9-]+$/.test(userId)) {
  throw new Error("Unexpected X user lookup response");
}
const configuration = $("Configuration").first().json;
const apiBase = configuration.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const token = data.next_token || "";
const base = apiBase + "/2/users/" + userId + "/tweets";
const params = "max_results=" + (demo ? "5" : "100") + "&post.fields=id,text,created_at,entities,attachments,note_post&expansions=referenced_posts";
const boundary = checkpoint.cursor ? "since_id=" + encodeURIComponent(checkpoint.cursor) : "start_time=" + encodeURIComponent(configuration.campaign_start_time);
const url = base + "?" + params + "&" + boundary + (token ? "&pagination_token=" + encodeURIComponent(token) : "");
return { user_id: userId, url: url, tweets: data.tweets || [], candidate_cursor: data.candidate_cursor || checkpoint.cursor,
  page_count: data.page_count || 0, tokens: data.tokens || [] };

这也开始暴露一些实现细节 :slight_smile: 我使用代理构建了工作流,并在过程中运行了一个“模拟 X API”以节省成本(因此那里有 localhost)。

使用流程中的“If 块”,只要还有更多页面,我们就可以继续遍历用户的时间线。X API 的一个很好的实现细节是,你只为实际获取的帖子付费,因此零结果查询不产生任何费用。

解析推文是使用这个小的脚本节点完成的:

const previous = $("Build page URL").item.json;
const response = $json;
if (!response || !response.meta || !Array.isArray(response.data || [])) {
  throw new Error("Incomplete X timeline response; refuse to update the post");
}
if (response.errors?.length) { throw new Error("X returned partial data/errors; refuse to update the post"); }
const next = response.meta.next_token || "";
if (next && (previous.tokens.includes(next) || previous.page_count >= 39)) {
  throw new Error("X pagination repeated or exceeded the 40-page safety limit");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
let candidate = previous.candidate_cursor;
for (const tweet of response.data || []) {
  if (typeof tweet.id !== "string" || !/^\d+$/.test(tweet.id)) { throw new Error("X post ID must be a decimal string"); }
  if (!candidate || greater(tweet.id, candidate)) { candidate = tweet.id; }
}
const normalize = post => {
  const note = post.note_post || post.note_tweet;
  return { ...post, text: note?.text || post.text, entities: note?.entities || post.entities,
    referenced_tweets: post.referenced_posts || post.referenced_tweets || [] };
};
const expanded = Object.fromEntries((response.includes?.posts || response.includes?.tweets || [])
  .map(post => [post.id, normalize(post)]));
const page = (response.data || []).map(normalize).map(post => ({ ...post,
  referenced_detail_urls: post.referenced_tweets.flatMap(ref => {
    const quote = expanded[ref.id];
    return quote ? [quote.text || "", ...(quote.entities?.urls || []).map(url => url.unwound_url || url.expanded_url || url.url)] : [];
  }) }));
const tweets = previous.tweets.concat(page);
return { user_id: previous.user_id, tweets: tweets, candidate_cursor: candidate, next_token: next,
  has_more: !!next, has_tweets: tweets.length > 0, page_count: previous.page_count + 1,
  tokens: next ? previous.tokens.concat(next) : previous.tokens };

添加智能

成功同步内容的关键成分是智能:



我们将推文列表作为输入传入 → 我们得到正确提取的 JSON 文档。这种类型的处理需要智能。

它既格式化内容,又找到最相关的信息片段。

更新 Discourse 帖子

我们使用一个脚本节点来按以下方式渲染表格:

const extracted = $("Poll result").first().json;
const checkpoint = $("Read checkpoint").first().json;
const apiBase = $("Configuration").first().json.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const raw = String($json.post?.raw || "");
const start = "<!-- x-ships-table:start -->";
const end = "<!-- x-ships-table:end -->";
const hasStart = raw.includes(start);
const hasEnd = raw.includes(end);
if (hasStart !== hasEnd ||
    (hasStart && (raw.indexOf(end) < raw.indexOf(start) ||
      raw.indexOf(start, raw.indexOf(start) + start.length) !== -1 ||
      raw.indexOf(end, raw.indexOf(end) + end.length) !== -1))) {
  throw new Error("Ambiguous workflow-owned table markers in designated post");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
if (checkpoint.cursor && extracted.candidate_cursor && greater(checkpoint.cursor, extracted.candidate_cursor)) {
  throw new Error("Refusing to move the timeline cursor backwards");
}
const bySlot = {};
for (const ship of checkpoint.ships.concat(extracted.ships)) {
  if (!Number.isInteger(ship.day) || ship.day < 1 || ship.day > 28 ||
      !new RegExp("^" + ship.day + "(?:\\.[1-9][0-9]?)?$").test(ship.slot) ||
      typeof ship.source_id !== "string" || !/^\d+$/.test(ship.source_id) ||
      typeof ship.announcement !== "string" || !ship.announcement || ship.announcement.length > 400 ||
      (ship.details_url && !/^https?:\/\/[^\s<>()[\]|\\]+$/.test(ship.details_url))) {
    throw new Error("Invalid stored ship slot");
  }
  if (!bySlot[ship.slot] || greater(ship.source_id, bySlot[ship.slot].source_id)) { bySlot[ship.slot] = ship; }
}
const escapeCell = value => String(value).replace(/[|\r\n]/g, " ").replace(/\\/g, "\\\\").replace(/([\[\]_*`])/g, "\\$1");
const rows = [];
for (let day = 1; day <= 28; day++) {
  const ships = Object.values(bySlot).filter(ship => ship.day === day)
    .sort((a, b) => a.slot.localeCompare(b.slot, undefined, { numeric: true }));
  if (!ships.length) { rows.push("| " + day + " | Pending | — | — | — |"); }
  else for (const ship of ships) {
    const provenance = demo
      ? "[Demo post](" + apiBase + "/demo/posts/" + encodeURIComponent(ship.source_id) + ") (synthetic)"
      : "[X post](https://x.com/" + encodeURIComponent($("Configuration").first().json.username) + "/status/" + encodeURIComponent(ship.source_id) + ")";
    const details = ship.details_url ? "[Details](" + ship.details_url + ")" : "—";
    rows.push("| " + day + " | " + escapeCell(ship.slot) + " | " + escapeCell(ship.announcement) + " | " + details + " | " + provenance + " |");
  }
}
const ships = Object.values(bySlot).sort((a, b) => a.day - b.day || a.slot.localeCompare(b.slot, undefined, { numeric: true }));
const updated = start + "\n| Day | Ship | Announcement | Details | Source |\n| --- | --- | --- | --- | --- |\n" + rows.join("\n") + "\n" + end;
const replacement = hasStart
  ? raw.slice(0, raw.indexOf(start)) + updated + raw.slice(raw.indexOf(end) + end.length)
  : raw + (raw ? (raw.endsWith("\n") ? "\n" : "\n\n") : "") + updated;
return { changed: replacement !== raw, raw: replacement, ships_json: JSON.stringify(ships),
  since_id: extracted.candidate_cursor, user_id: extracted.user_id };

小窍门在于 Discourse 帖子中有一个特殊的标记:


这样,在生成表格后,我们可以在更新表格之前检查是否有任何变化。

如果表格发生变化,我们更新跟踪数据表:

我是如何构建这个的?

工作流是极其强大的功能,手动构建像这样的大型工作流将花费很长时间。

我没有那样做,而是依靠代理。我在所有开发中使用 term-llm.com,它非常干净地集成到 dv 中。当然,你可以使用你选择的任何代理,但这里的一个关键功能是拥有一个可用的 Discourse 环境。

我的步骤是:

  1. dv new workflow-build - 创建一个干净的新环境。

    1. 我使用了一个特殊的 dv hook,它启动一个 term-llm serve web 服务,该服务连接回我的 hub,这为我提供了一个即时的“丰富代理”,它在一个我可以随意运行的临时环境中。
  2. 我提示它:


我在构建过程中使用了 GPT 6.1 Sol,并在一次尝试后就已经有了初步成果:

  1. 我对其进行了优化
  • 第一次尝试没有使用数据表 - 我提示将存储移动到那里

  • 接下来,我迭代将所有配置移到一个中心节点,以便更容易管理

  1. 我导出了工作流 - 然后导入到生产环境

  2. 我与代理进行了一些迭代,因为在生产环境中有一些缺失的边缘情况(我们的导出 xml 不包括数据表定义和代理定义)

如果你需要构建类似的东西怎么办?

这篇帖子将对代理非常有帮助,一个简单的提示,例如:

将账户 A、B、C 上关于“X”的重要讨论同步到我的主题中,遵循 META 链接中的类似方法

这将允许你从 X 进行任意数据同步。

社交 → Discourse 是社区的一个非常重要的向量,工作流允许你保持一个长期运行的记录,开放更细致的讨论,而无需强迫公司中现有的人员将所有内容发布集中到 Discourse 上。

工作流是你可用于这种同步的构建块,这里的选项非常广泛。

  • 当特定人员在社交上发布内容时,发布影子主题
  • 根据社交上的活动保持长期主题的最新状态
  • … 以及更多
1 个赞