Capture first, sync second
A headless SitecoreAI site we run collects donations and about a dozen supporter forms: inquiries, document requests, cancellations, detail updates. The CRM is Salesforce, and every one of those submissions has to land there reliably, attributed to the right campaign.
Instead of a third-party integration service, we built a small hub we own: a Next.js app that captures every submission into Postgres first, syncs it to Salesforce second, and exposes an admin dashboard that runs inside SitecoreAI as a Sitecore Marketplace app. This post covers the three pieces that carry the weight: the auth, the sync loop, and the embedded dashboard.
The pipeline, in order
The donor's browser never talks to Salesforce, and neither does anything else on the public path. The head app proxies form submissions server-side to the hub's API. The hub validates the payload, writes the record to Postgres, and returns. The Salesforce sync happens after the fact: fired in the background immediately after a successful capture, with a scheduled batch sweep behind it as the safety net.
browser ──▶ head app /api/* ──▶ hub API ──▶ Postgres (capture, source of record)
│
└──▶ Salesforce sync (async + batch sweep)
The ordering is the entire point. A Salesforce outage, an expired certificate, or a bad field mapping cannot lose a donation, because the donation exists in Postgres before any CRM call is attempted. Postgres is the source of record for web capture; Salesforce remains the source of truth for CRM data. Every synced table carries the same three columns so the state is visible and sweepable:
sfId String?
sfSynced Boolean @default(false)
sfSyncedAt DateTime?
The batch sweep is then a findMany over rows where sfSynced is false.
Server-to-server auth: the JWT Bearer flow without an SDK
We skipped the Salesforce client libraries. The OAuth 2.0 JWT Bearer flow wants an RS256-signed JWT whose iss is the connected app's consumer key, sub is the integration user, aud is the login server, and exp is a Unix timestamp [1]. Posting that assertion to the token endpoint returns an access token and the org's instance URL. The flow never issues a refresh token, so each sync simply authenticates again. A jti claim is optional, but if you send one Salesforce rejects any repeat, which closes the replay window at no cost [1]. That is about a hundred lines with Node's built-in crypto and fetch, and owning those lines made every timeout and error path explicit:
import crypto from 'node:crypto';
type SalesforceAuth = { accessToken: string; instanceUrl: string };
function buildJWT(): string {
const now = Math.floor(Date.now() / 1000);
const header = { alg: 'RS256', typ: 'JWT' };
const payload = {
iss: process.env.SF_CLIENT_ID,
sub: process.env.SF_USERNAME,
aud: process.env.SF_LOGIN_URL,
exp: now + 300,
jti: crypto.randomUUID(),
};
const b64url = (obj: object) => Buffer.from(JSON.stringify(obj)).toString('base64url');
const signingInput = `${b64url(header)}.${b64url(payload)}`;
const signer = crypto.createSign('RSA-SHA256');
signer.update(signingInput);
return `${signingInput}.${signer.sign(getPrivateKey(), 'base64url')}`;
}
export async function authenticate(): Promise {
const response = await fetch(`${process.env.SF_LOGIN_URL}/services/oauth2/token`, {
method: 'POST',
headers: { 'Content-Type': 'application/x-www-form-urlencoded' },
body: new URLSearchParams({
grant_type: 'urn:ietf:params:oauth:grant-type:jwt-bearer',
assertion: buildJWT(),
}),
signal: AbortSignal.timeout(10_000),
});
const text = await response.text();
let data: Record;
try {
data = JSON.parse(text);
} catch {
throw new Error(`Salesforce auth failed: HTTP ${response.status}: ${text.slice(0, 200)}`);
}
if (!response.ok) {
throw new Error(`Salesforce auth failed: ${data.error} - ${data.error_description}`);
}
return { accessToken: data.access_token, instanceUrl: data.instance_url };
}
Three production notes that cost us time.
PEM keys and environment variables don't mix cleanly. Environment variables flatten the private key onto one line. Node has shipped OpenSSL 3 since version 17, and OpenSSL 3 tightened what it accepts [4]; in our case a flattened key failed to decode. Don't try to preserve the original formatting. Normalize unconditionally:
function getPrivateKey(): string {
const raw = process.env.SF_PRIVATE_KEY!.replace(/
/g, '
');
const isPkcs1 = /BEGIN RSA PRIVATE KEY/.test(raw);
const base64 = raw
.replace(/-----(BEGIN|END)\s+(RSA\s+)?PRIVATE\s+KEY-----/g, '')
.replace(/\s+/g, '');
const label = isPkcs1 ? 'RSA PRIVATE KEY' : 'PRIVATE KEY';
return [
`-----BEGIN ${label}-----`,
...(base64.match(/.{1,64}/g) ?? []),
`-----END ${label}-----`,
].join('
');
}
Time out every call. AbortSignal.timeout(), available since Node 17.3 [5], goes on every fetch. The rejected promise carries a TimeoutError, which our retry classifier treats as transient. In our experience Salesforce answers in well under two seconds. On serverless, one hanging call otherwise consumes the entire function budget of whatever triggered the sync.
Parse defensively. When the login endpoint has a bad day, the 5xx comes back as plain text from a gateway, not JSON. Read the body as text first, then parse. An auth-failure message that includes the status code beats a JSON.parse stack trace.
A sync loop with retries that are safe to retry
The background sync is fire-and-forget from the caller's perspective, but internally it retries, carefully. The first rule of automatic retries: make them safe before making them automatic. Three properties do that here. The create path dedupes by an external transaction ID before inserting, so a re-run cannot double-create. The update path is a PATCH of the same values, which is idempotent. And auth is stateless. Only with those in place does the retry loop earn its keep.
If you are designing the Salesforce side from scratch, mark that transaction ID as an External ID field and use the upsert resource: a PATCH to /sobjects/<Object>/<ExternalIdField>/<value>/ creates or updates in one call, and answers with HTTP 300 if the value matches more than one record [3]. We inherited an object without one, so our create path queries by transaction ID first and adopts the existing record if it finds one.
The second rule: not every error deserves a retry. A gateway timeout heals in seconds. A 400 INVALID_FIELD is permanent, and retrying it delays the caller and floods the logs. Salesforce reports API-limit exhaustion as 403 REQUEST_LIMIT_EXCEEDED rather than as a 429 [2]. That is an org-wide quota, not a transient condition, so retrying a few seconds later will not help either, and the classifier below leaves it to the batch sweep:
type SyncResult = { success: boolean; sfId?: string; error?: string };
const RETRY_DELAYS_MS = [2_000, 5_000];
function isRetryableSyncError(error: string): boolean {
if (error.includes('not found')) return false;
const status = error.match(/\b(4\d\d)\b/)?.[1];
return !status || status === '429';
}
export async function syncDonation(donationId: string): Promise {
let last: SyncResult = { success: false, error: 'not attempted' };
for (let attempt = 0; attempt <= RETRY_DELAYS_MS.length; attempt++) {
if (attempt > 0) await sleep(RETRY_DELAYS_MS[attempt - 1]);
try {
last = await syncOnce(donationId);
} catch (err) {
last = { success: false, error: err instanceof Error ? err.message : String(err) };
await logToErrorTable(donationId, 'SYNC_UNEXPECTED', last.error);
}
if (last.success || !isRetryableSyncError(last.error ?? '')) return last;
}
return last;
}
One honest caveat. Classifying on the error message is the weakest part of this design: any standalone three-digit number starting with 4 in a message could be mistaken for a status code. In a new codebase, have the client throw an error object that carries the HTTP status and classify on that field. We kept the string check because it matched the client we already had, and the messages it produces are under our control.
Two details matter more than they look.
Every failed attempt is written to an error-log table with the request payload, not only to the console. That payload contains personal data, so the table gets the same access controls and retention policy as the records themselves. Console logs are for developers. The error table feeds the admin dashboard, where a person can see the failure and click re-sync.
The batch sweep calls the same function with retries disabled. The sweep is the retry mechanism at that layer, and stacking retry loops multiplies delays.
Before this hardening, a Salesforce login timeout at payment-callback time left the record sitting "Not Synced" until someone noticed and clicked re-sync in the dashboard. After it, transient blips heal themselves and only genuine mapping problems need a human.
One last practice: reconcile end-to-end on a schedule. Pull the ID list from Salesforce and diff it against Postgres, record by record. Sync-state columns on your side can lie, because a process can die between the CRM write and the flag update. The external transaction ID is the keystone of the whole design. It is what makes dedup, idempotent retries, and reconciliation possible. If you adopt one thing from this post, put a unique external ID on your CRM object.
Operations inside SitecoreAI: the Marketplace app
The capture-and-sync pipeline would work headless, but pipelines without visibility rot. The hub's admin dashboard shows submissions, sync status, and the error-log review queue. It also lets the campaign team set up campaigns whose attribution codes come straight from Salesforce: the campaign wizard searches Salesforce campaigns live, with a SOQL query through the same JWT auth, so codes are selected, never re-typed.
What makes it feel native is the Sitecore Marketplace. The hub is registered as a Marketplace app, so SitecoreAI renders it in an iframe inside the Sitecore shell, alongside Pages and the other editing tools, with the Marketplace SDK connecting the two over the browser's postMessage API [6][7]. Identity comes from the Sitecore session via the SDK and is exchanged for a JWT session on our side. There is no second login and no local user table, and every admin action lands in an activity log attributed to the person's Sitecore identity. The embedding itself is covered in more depth in Replacing xDB on SitecoreAI with a Marketplace App.
For the content and campaign team, the mental model collapses to one sentence: the website's data lives in the same place the website is edited. Nobody has to know that a separate app, a Postgres database, and a sync pipeline sit behind that tab.
Two operational footnotes from running it.
Grant the download permission at registration. Iframe-embedded apps cannot trigger file downloads until the app registration allows it, so CSV exports silently fail.
Keep a break-glass login. Ours is IP-allowlisted and off by default, for the day the embedding itself has a problem.
What we'd tell you to copy
- Write locally first, sync second. A CRM outage should cost you freshness, never data.
- Idempotency before automation. External IDs and dedup-before-create make retries safe.
- Classify errors. Retry a 504, fail a 400 fast, and leave a request-limit 403 to the sweep.
- Time out every upstream call on serverless. One hung call starves the function that made it.
- Put sync state in front of non-developers. The dashboard answers "did it reach the CRM?", not a log search.
- Reconcile on a schedule. Your flags can lie; the other system is the check.
Sources
- OAuth 2.0 JWT Bearer Flow for Server-to-Server Integration — Salesforce Help
- Status Codes and Error Responses — Salesforce REST API Developer Guide
- Insert or Update (Upsert) a Record Using an External ID — Salesforce REST API Developer Guide
- Node v17.0.0 (Current) — Node.js Blog
- AbortSignal.timeout(delay) — Node.js Documentation
- Marketplace SDK — Sitecore Documentation
- Sitecore/marketplace-sdk — GitHub