Aggiornamento automatico di un post su Discourse in base a un feed di X

Ho appena configurato il seguente su OpenAI

Come scrive Tibo su X, Discourse ora aggiorna automaticamente un post che tiene traccia dei tweet.

Come funziona?

Una tabella dati come cache

Discourse Workflows include una funzionalità di tabella dati.

Ciò ci consente di memorizzare informazioni strutturate tra un’esecuzione e l’altra del workflow.

Un elemento cruciale nell’integrazione con X è evitare il superamento dei costi dell’API. Per farlo, memorizziamo una cache di tutte le informazioni estratte per evitare di interrogare nuovamente i tweet vecchi.

Una tabella dati è perfetta per questo scopo.

Un nodo di configurazione

“Imposta campi” è perfetto per i nodi di configurazione. In questo caso avevo bisogno di alcune configurazioni “per workflow” utilizzate in tutto il flusso, come l’ID del topic in cui pubblicare, l’ora di inizio della campagna, i nomi utente e così via.

Aiuta a mantenere il workflow più facile da comprendere.

Richieste HTTP a X

Il workflow si basa su 2 endpoint di X:

Uno molto semplice per cercare l’ID di X di Tibo.

Un secondo per cercare la timeline:

const data = $json;
const checkpoint = $("Read checkpoint").first().json;
const userId = data.user_id || data.data?.id;
if (!userId || !/^[A-Za-z0-9-]+$/.test(userId)) {
  throw new Error("Unexpected X user lookup response");
}
const configuration = $("Configuration").first().json;
const apiBase = configuration.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const token = data.next_token || "";
const base = apiBase + "/2/users/" + userId + "/tweets";
const params = "max_results=" + (demo ? "5" : "100") + "&post.fields=id,text,created_at,entities,attachments,note_post&expansions=referenced_posts";
const boundary = checkpoint.cursor ? "since_id=" + encodeURIComponent(checkpoint.cursor) : "start_time=" + encodeURIComponent(configuration.campaign_start_time);
const url = base + "?" + params + "&" + boundary + (token ? "&pagination_token=" + encodeURIComponent(token) : "");
return { user_id: userId, url: url, tweets: data.tweets || [], candidate_cursor: data.candidate_cursor || checkpoint.cursor,
  page_count: data.page_count || 0, tokens: data.tokens || [] };

Questo inizia anche a rivelare alcuni dettagli implementativi :slight_smile: Ho costruito il workflow utilizzando un agente e ho eseguito una “falsa API di X” durante il processo per risparmiare sui costi (ecco perché c’è localhost lì)

Utilizzando i blocchi “If” del flusso, possiamo continuare a iterare sulla timeline di un utente finché ci sono altre pagine. Un bel dettaglio implementativo dell’API di X è che si paga solo per i post effettivi che si ottengono, quindi le query con zero risultati non costano nulla.

L’analisi dei tweet viene eseguita utilizzando questo piccolo nodo script:

const previous = $("Build page URL").item.json;
const response = $json;
if (!response || !response.meta || !Array.isArray(response.data || [])) {
  throw new Error("Incomplete X timeline response; refuse to update the post");
}
if (response.errors?.length) { throw new Error("X returned partial data/errors; refuse to update the post"); }
const next = response.meta.next_token || "";
if (next && (previous.tokens.includes(next) || previous.page_count >= 39)) {
  throw new Error("X pagination repeated or exceeded the 40-page safety limit");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
let candidate = previous.candidate_cursor;
for (const tweet of response.data || []) {
  if (typeof tweet.id !== "string" || !/^\d+$/.test(tweet.id)) { throw new Error("X post ID must be a decimal string"); }
  if (!candidate || greater(tweet.id, candidate)) { candidate = tweet.id; }
}
const normalize = post => {
  const note = post.note_post || post.note_tweet;
  return { ...post, text: note?.text || post.text, entities: note?.entities || post.entities,
    referenced_tweets: post.referenced_posts || post.referenced_tweets || [] };
};
const expanded = Object.fromEntries((response.includes?.posts || response.includes?.tweets || [])
  .map(post => [post.id, normalize(post)]));
const page = (response.data || []).map(normalize).map(post => ({ ...post,
  referenced_detail_urls: post.referenced_tweets.flatMap(ref => {
    const quote = expanded[ref.id];
    return quote ? [quote.text || "", ...(quote.entities?.urls || []).map(url => url.unwound_url || url.expanded_url || url.url)] : [];
  }) }));
const tweets = previous.tweets.concat(page);
return { user_id: previous.user_id, tweets: tweets, candidate_cursor: candidate, next_token: next,
  has_more: !!next, has_tweets: tweets.length > 0, page_count: previous.page_count + 1,
  tokens: next ? previous.tokens.concat(next) : previous.tokens };

Aggiunta di intelligenza

L’ingrediente chiave per sincronizzare con successo il contenuto è l’intelligenza:



Passiamo un elenco di tweet come input → otteniamo un documento JSON correttamente estratto. Questo tipo di processo richiede intelligenza.

Formatta il contenuto E trova i pezzi di informazione più rilevanti.

Aggiornamento del post di Discourse

Utilizziamo un nodo script per rendere la tabella per:

const extracted = $("Poll result").first().json;
const checkpoint = $("Read checkpoint").first().json;
const apiBase = $("Configuration").first().json.api_base_url.replace(/\/$/, "");
const demo = /^http:\/\/(?:localhost|127\.0\.0\.1):[1-9]\d*$/.test(apiBase);
const raw = String($json.post?.raw || "");
const start = "<!-- x-ships-table:start -->";
const end = "<!-- x-ships-table:end -->";
const hasStart = raw.includes(start);
const hasEnd = raw.includes(end);
if (hasStart !== hasEnd ||
    (hasStart && (raw.indexOf(end) < raw.indexOf(start) ||
      raw.indexOf(start, raw.indexOf(start) + start.length) !== -1 ||
      raw.indexOf(end, raw.indexOf(end) + end.length) !== -1))) {
  throw new Error("Ambiguous workflow-owned table markers in designated post");
}
const greater = (a, b) => a.length > b.length || (a.length === b.length && a > b);
if (checkpoint.cursor && extracted.candidate_cursor && greater(checkpoint.cursor, extracted.candidate_cursor)) {
  throw new Error("Refusing to move the timeline cursor backwards");
}
const bySlot = {};
for (const ship of checkpoint.ships.concat(extracted.ships)) {
  if (!Number.isInteger(ship.day) || ship.day < 1 || ship.day > 28 ||
      !new RegExp("^" + ship.day + "(?:\\.[1-9][0-9]?)?$").test(ship.slot) ||
      typeof ship.source_id !== "string" || !/^\d+$/.test(ship.source_id) ||
      typeof ship.announcement !== "string" || !ship.announcement || ship.announcement.length > 400 ||
      (ship.details_url && !/^https?:\/\/[^\s<>()[\]|\\]+$/.test(ship.details_url))) {
    throw new Error("Invalid stored ship slot");
  }
  if (!bySlot[ship.slot] || greater(ship.source_id, bySlot[ship.slot].source_id)) { bySlot[ship.slot] = ship; }
}
const escapeCell = value => String(value).replace(/[|\r\n]/g, " ").replace(/\\/g, "\\\\").replace(/([\[\]_*`])/g, "\\$1");
const rows = [];
for (let day = 1; day <= 28; day++) {
  const ships = Object.values(bySlot).filter(ship => ship.day === day)
    .sort((a, b) => a.slot.localeCompare(b.slot, undefined, { numeric: true }));
  if (!ships.length) { rows.push("| " + day + " | Pending | — | — | — |"); }
  else for (const ship of ships) {
    const provenance = demo
      ? "[Demo post](" + apiBase + "/demo/posts/" + encodeURIComponent(ship.source_id) + ") (synthetic)"
      : "[X post](https://x.com/" + encodeURIComponent($("Configuration").first().json.username) + "/status/" + encodeURIComponent(ship.source_id) + ")";
    const details = ship.details_url ? "[Details](" + ship.details_url + ")" : "—";
    rows.push("| " + day + " | " + escapeCell(ship.slot) + " | " + escapeCell(ship.announcement) + " | " + details + " | " + provenance + " |");
  }
}
const ships = Object.values(bySlot).sort((a, b) => a.day - b.day || a.slot.localeCompare(b.slot, undefined, { numeric: true }));
const updated = start + "\n| Day | Ship | Announcement | Details | Source |\n| --- | --- | --- | --- | --- |\n" + rows.join("\n") + "\n" + end;
const replacement = hasStart
  ? raw.slice(0, raw.indexOf(start)) + updated + raw.slice(raw.indexOf(end) + end.length)
  : raw + (raw ? (raw.endsWith("\n") ? "\n" : "\n\n") : "") + updated;
return { changed: replacement !== raw, raw: replacement, ships_json: JSON.stringify(ships),
  since_id: extracted.candidate_cursor, user_id: extracted.user_id };

Il piccolo trucco è che il post di Discourse ha un marcatore speciale:


In questo modo, dopo aver generato la tabella, possiamo verificare se qualcosa è cambiato prima di aggiornare la tabella.

Se la tabella cambia, aggiorniamo la tabella dati di tracciamento:

Come l’ho costruito?

I workflow sono funzionalità incredibilmente potenti; costruirne di grandi come questo a mano richiederebbe molto tempo.

Invece di farlo, mi sono affidato a un agente. Uso term-llm.com per tutto il mio sviluppo, si integra molto bene con dv. Naturalmente puoi usare l’agente che preferisci, ma una funzione chiave qui è avere un ambiente Discourse funzionante.

I miei passaggi sono stati:

  1. dv new workflow-build - per creare un nuovo ambiente pulito.

    1. Uso un hook dv speciale che avvia un servizio web term-llm serve che si riconnette al mio hub, il che mi dà un “agente ricco” istantaneo in un ambiente usa e getta in cui posso fare ciò che voglio.
  2. Gli ho dato un prompt:


Ho usato GPT 6.1 Sol per la costruzione e avevo già qualcosa di funzionante dopo un solo tentativo:

  1. L’ho rifinito
  • Il primo tentativo non utilizzava una tabella dati - ho dato un prompt per spostare lo storage lì

  • Successivamente ho iterato per spostare tutta la configurazione in un nodo centrale per renderla più facile da gestire

  1. Ho esportato il workflow - e poi l’ho importato in produzione

  2. Ho iterato un po’ con l’agente perché in produzione c’erano alcuni casi limite mancanti (il nostro xml di esportazione non include le definizioni delle tabelle dati e delle definizioni degli agenti)

E se dovessi costruire qualcosa del genere?

Questo post sarà incredibilmente utile per gli agenti, un prompt semplice lungo le linee di:

Sincronizza le discussioni importanti su “X” dagli account A, B, C nel mio topic seguendo un metodo simile in LINK TO META

Ti consentirà di eseguire una sincronizzazione dati arbitraria da X.

Social → Discourse è un vettore molto importante per la community, i workflow ti consentono di mantenere un registro a lungo termine che è aperto a discussioni più sfumate senza costringere le persone esistenti nella tua azienda a centralizzare tutti i post di contenuto su Discourse.

I workflow sono il mattone che puoi usare per questo tipo di sincronizzazione e le opzioni qui sono piuttosto ampie.

  • Pubblica topic ombra quando persone specifiche pubblicano contenuti sui social
  • Mantieni aggiornato un topic a lungo termine in base all’attività sui social
  • … e molto altro
1 Mi Piace