Skip to content

feat(webapp): hosted webhook ingress, delivery pipeline, and dashboard - #4344

Open
ericallam wants to merge 32 commits into
mainfrom
feat/hosted-webhook-ingress
Open

feat(webapp): hosted webhook ingress, delivery pipeline, and dashboard#4344
ericallam wants to merge 32 commits into
mainfrom
feat/hosted-webhook-ingress

Conversation

@ericallam

@ericallam ericallam commented Jul 23, 2026

Copy link
Copy Markdown
Member

Summary

The server half of hosted webhooks: the public ingress endpoint, signature verification, the delivery pipeline (Postgres partitioned storage + ClickHouse for ordering), the in-app partition manager, the HTTP API, and the dashboard (Deliveries, Endpoints, and the in-app test console).

The public SDK and docs half is #4537. That PR carries the user-facing API (webhook(), chat.event / chat.channels, the @trigger.dev/slack connector) and builds on the shared @trigger.dev/core schemas that ship here.

Shipping behind a flag

A WEBHOOK_ENABLED env var (default off) gates the public ingress route and the engine worker plus partition cron, so merging and deploying this changes nothing in production until it is flipped on per environment. The dashboard is separately gated per org by the hasWebhooksAccess feature flag.

Note on packages

This PR includes the @trigger.dev/core schema additions the server compiles against, but carries no changeset. Core is not consumed independently of the SDK, so it is released together with the SDK via #4537. Keeping its changeset off main means no release cut from main publishes it early.

@changeset-bot

changeset-bot Bot commented Jul 23, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: ead2f98

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@coderabbitai

coderabbitai Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

This PR introduces a hosted webhooks platform spanning core schemas/types, a new internal webhook-engine package (signature verification, filtering, signing, delivery routing, partitioning), a webhook-sources provider registry with sample payloads, database/ClickHouse/replication infrastructure, webapp persistence, services, presenters, public API routes, and dashboard UI for managing webhook endpoints and deliveries. It also adds a @trigger.dev/slack channel connector package, extends the trigger-sdk chat runtime with webhook events/channels support, updates CLI worker manifests, and adds extensive documentation for the new webhooks feature.

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 20.31% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Description check ⚠️ Warning The description explains the implementation and rollout, but it omits the required issue link, checklist, testing, changelog, and screenshots sections. Add the template sections, include the related issue, record completed checklist items, document testing steps, summarize the changelog, and add screenshots or state that none apply.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main hosted webhook ingress, delivery pipeline, and dashboard changes.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/hosted-webhook-ingress

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

github-advanced-security[bot]

This comment was marked as resolved.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Due to the large number of review comments, Critical severity comments were prioritized as inline comments.

🟠 Major comments (23)
packages/cli-v3/src/dev/devSupervisor.ts-476-476 (1)

476-476: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Fail dequeued runs whose worker is unavailable.

Removing #failRunWithMissingWorker leaves this attempt without a controller or terminal API update. It can remain stuck or be repeatedly redelivered. Restore the terminal failure call before continuing.

.changeset/hosted-webhook-ingress.md-8-14 (1)

8-14: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Do not publish this as generally available while the feature flag remains disabled.

The PR objective says hosted webhooks are unreleased and disabled by default, but these changes publish a release and public documentation that present it as available.

  • .changeset/hosted-webhook-ingress.md#L8-L14: defer the public release until rollout, or use the project’s preview-release mechanism.
  • docs/docs.json#L164-L176: keep the navigation out of public docs until the feature is enabled.
  • docs/webhooks/overview.mdx#L7-L9: clearly mark and gate the feature as preview if these docs must ship first.
internal-packages/clickhouse/src/webhookDeliveries.ts-140-154 (1)

140-154: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Exclude tombstoned deliveries from grouped counts.

count(DISTINCT delivery_id) deduplicates status versions but still counts a delivery whose latest row has _is_deleted = 1. Unlike the list and total-count builders, this query skips FINAL, so endpoint totals can be stale after deletes. Use FINAL or a version-aware argMax aggregation.

Proposed fix
-      "SELECT webhook_endpoint_id, count(DISTINCT delivery_id) AS count FROM trigger_dev.webhook_deliveries_v1",
+      "SELECT webhook_endpoint_id, count() AS count FROM trigger_dev.webhook_deliveries_v1 FINAL",
internal-packages/replication/src/client.ts-450-462 (1)

450-462: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Validate existing publications against publish_via_partition_root.

#validatePublicationConfiguration() checks the publication table and actions, but not pg_publication.pubviaroot. A partitioned source publication created without pubviaroot = true will be reused, so replication will stream events using child-partition identity/schema instead of the intended root-table semantics. Compare pubviaroot when publishViaPartitionRoot is enabled and fail with remediation guidance such as ALTER PUBLICATION ... SET (publish_via_partition_root = true).

internal-packages/webhook-sources/src/index.ts-22-27 (1)

22-27: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Make individual samples uniquely addressable.

getSample() returns the first duplicate (provider, eventType). Selecting “Inbound image message” returns the text-message body; Zendesk’s later ticket.updated examples are similarly unreachable. Add a stable sample ID to manifest items and use it for lookup.

Based on supplied context, apps/webapp/app/routes/resources.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks.samples.ts:53-57 calls this function with only provider and event type.

internal-packages/webhook-sources/catalog/providers.json-1-1 (1)

1-1: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Catalog checklist contradicts the registry files' own "sample-only" doc comments for workos and zendesk.

Both providers.json entries claim roundTrip/producer are done and tier first-class, but the corresponding registry files explicitly state, in the present tense, that the provider currently ships sample-only pending a custom verifier config — the opposite conclusion. Compare with anthropic/hubspot/twilio in the same file, which correctly downgraded tier and checklist when a similar caveat applied.

  • internal-packages/webhook-sources/catalog/providers.json#L903-923: either downgrade the workos entry to tier: "sample-only" with only registryEntry/samples checked, or confirm a custom verifier + round-trip test genuinely exists and update registry/workos.ts's stale comment.
  • internal-packages/webhook-sources/catalog/providers.json#L660-681: same reconciliation needed for the zendesk entry.
  • internal-packages/webhook-sources/src/registry/workos.ts#L4-6: if the checklist is correct, update this comment — it currently reads as if verification is still pending.
  • internal-packages/webhook-sources/src/registry/zendesk.ts#L4-8: same — update this comment if verification is in fact complete.
internal-packages/webhook-sources/src/handAuthored/sentry.ts-58-60 (1)

58-60: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Keep Sentry-Hook-Resource resource-level.

All four samples set the resource header to <resource>.<action>, but Sentry delivers only the resource in the reserved Sentry-Hook-Resource header (issue or error here), while the action lives in the payload. Update issue.created/resolved/assigned to issue and error.created to error; also apply the same change to internal-packages/webhook-sources/src/handAuthored/sentry.ts:114-116, 174-176, 253-255.

Source: MCP tools

internal-packages/webhook-sources/src/registry/telegram.ts-17-18 (1)

17-18: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Do not use the Telegram message object as the event-type path.

Line 17 resolves to an object, not an event-type value, and is absent for most Telegram Update variants. Extend EventTypeSource and its consumer to support top-level-key discrimination, or provide a provider-specific extractor before enabling this entry.

internal-packages/webhook-sources/src/registry/whatsapp.ts-18-18 (1)

18-18: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Do not use field as WhatsApp’s event discriminant.

The configured path is "messages" for both inbound messages and status updates, so routing and filtering cannot distinguish them. Use an extractor that checks value.messages versus value.statuses, or extend the source contract to support key-presence discrimination.

apps/webapp/app/db.server.ts-270-278 (1)

270-278: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Unhandled promise rejection risk on client.$connect().

client.$connect() at line 275 is fire-and-forget: no await, no .catch(). This function is now the single construction point for 4 Prisma clients (main writer/reader plus the new webhook writer/reader). If any of DATABASE_URL, DATABASE_READ_REPLICA_URL, WEBHOOK_DATABASE_URL, or WEBHOOK_DATABASE_READ_REPLICA_URL is misconfigured, the rejected connect promise goes unhandled — in modern Node this crashes the process instead of surfacing a clear, logged error. The "connected" console.log right after also prints unconditionally, independent of whether the connection actually succeeded.

🛡️ Proposed fix: observe the connect promise
   // connect eagerly
-  client.$connect();
-
-  console.log(`🔌 ${clientType} prisma client connected`);
+  client
+    .$connect()
+    .then(() => console.log(`🔌 ${clientType} prisma client connected`))
+    .catch((error) => {
+      logger.error(`Failed to connect ${clientType} prisma client`, {
+        clientType,
+        error: error instanceof Error ? error.message : String(error),
+      });
+    });
apps/webapp/app/services/realtime/sessions.server.ts-118-125 (1)

118-125: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Use findFirst instead of findUnique.

findSessionByExternalId queries with findUnique. The composite key works equally with findFirst, which is the required convention in this codebase.

♻️ Proposed change
-  return prisma.session.findUnique({
+  return prisma.session.findFirst({
     where: { runtimeEnvironmentId_externalId: { runtimeEnvironmentId: environment.id, externalId } },
   });

As per path instructions: "Always use Prisma findFirst instead of findUnique."

Source: Path instructions

apps/webapp/app/env.server.ts-1523-1524 (1)

1523-1524: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Default the webhook feature flags to opt-in.

entry.server.tsx touches the webhookEngine singleton at startup, so the worker starts unless its disabled flag is "0". Since WEBHOOK_WORKER_ENABLED defaults from WORKER_ENABLED (?? "true"), existing deploys with WORKER_ENABLED=true will start the webhook redis-worker on upgrade. WEBHOOK_INGRESS_ENABLED also defaults to "1". Hard-default both flags to "0" unless each feature is behind a separate rollout gate.

Source: Learnings

apps/webapp/app/routes/resources.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks.deliveries.live.ts-65-77 (1)

65-77: 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Auth/scope check is skipped when no filters are supplied.

The empty-result short-circuit (line 73-75) runs before loadProjectEnvironmentFromRequest (line 77), so a request to this route with no deliveryIds/since params never authenticates or validates org/project/env scope — it just returns 200 { deliveries: [] }. Move the check after resolving project/environment so the route consistently enforces auth regardless of which filters are present.

🔒 Proposed fix: authenticate/scope before short-circuiting
   const newDeliveriesSince =
     includeNewDeliveries && since !== undefined ? since : undefined;

-  if (deliveryIds.length === 0 && newDeliveriesSince === undefined) {
-    return typedjson({ deliveries: [] });
-  }
-
   const { project, environment } = await loadProjectEnvironmentFromRequest(request, params);
 
+  if (deliveryIds.length === 0 && newDeliveriesSince === undefined) {
+    return typedjson({ deliveries: [] });
+  }
+
   const clickhouse = await clickhouseFactory.getClickhouseForOrganization(
apps/webapp/app/v3/services/createDeploymentBackgroundWorkerV4.server.ts-230-246 (1)

230-246: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Non-ServiceValidationError webhook sync failures are silently swallowed.

Unlike the schedules block just above (lines 199-228), a transient/unexpected webhooksError that isn't a ServiceValidationError is only logged — the deployment then proceeds to DEPLOYING as if webhook sync succeeded, leaving WebhookEndpoint rows stale/missing. createDeploymentBackgroundWorkerV3.server.ts doesn't have this gap since its surrounding try/catch fails the deployment on any error.

🐛 Proposed fix: fail the deployment on any webhook sync error
       if (webhooksError) {
         logger.error("Error syncing declarative webhooks", { error: webhooksError });
         if (webhooksError instanceof ServiceValidationError) {
           await this.#failBackgroundWorkerDeployment(deployment, webhooksError);
           throw webhooksError;
         }
+
+        const serviceError = new ServiceValidationError("Error syncing declarative webhooks");
+        await this.#failBackgroundWorkerDeployment(deployment, serviceError);
+        throw serviceError;
       }
apps/webapp/app/routes/api.v1.webhooks.endpoints.$endpointId.disable.ts-26-33 (1)

26-33: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Read-after-write via replica right after a primary write can return stale status.

webhookPrisma.webhookEndpoint.update writes to the primary, but findWebhookEndpointResource reads through ApiWebhookEndpointPresenter, which queries webhookReplica. If replication lags, the API response to this "disable" call can still show the endpoint's old status instead of inactive.

🛡️ Proposed fix
-    await webhookPrisma.webhookEndpoint.update({
+    const updated = await webhookPrisma.webhookEndpoint.update({
       where: { id: endpoint.id },
       data: { status: "INACTIVE" },
     });
     webhookEngine.invalidateEndpoint(endpoint.opaqueId);

-    return json(await findWebhookEndpointResource(authentication, params.endpointId));
+    // Build the response from the primary-write result to avoid replica-lag staleness.
+    return json(toApiEndpointFromPrimary(updated));

Alternatively, have the presenter accept an explicit Prisma client so this route can pass webhookPrisma for a guaranteed-fresh read.

apps/webapp/app/presenters/v3/ApiWebhookDeliveryPresenter.server.ts-92-105 (1)

92-105: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Pagination cursor/direction disagree when both page[after] and page[before] are supplied.

cursor resolves page[after] first, but direction resolves page[before] first. If both are present, the request is sent as direction: "backward" while using the page[after] value as the cursor — an inconsistent pair. Per the established convention, page[before] should win for both the cursor value and the direction.

🐛 Proposed fix
-        cursor: searchParams["page[after]"] ?? searchParams["page[before]"],
+        cursor: searchParams["page[before]"] ?? searchParams["page[after]"],
         direction: searchParams["page[before]"] ? "backward" : "forward",

Based on learnings, "it is an established shared convention to allow both cursor query params page[after] and page[before]... When both are present, page[before] must take precedence (i.e., it should be used/wins)."

Source: Learnings

apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks.deliveries.$deliveryParam/route.tsx-134-175 (1)

134-175: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Fix the hook call order before the early return.

useState is currently reached only after delivery is truthy, so any render transition between the missing-delivery branch and the delivery branch changes the hook order/instance count within the same component instance. Move this hook above the !delivery return.

apps/webapp/app/routes/api.v1.webhooks.endpoints.$endpointId.enable.ts-26-33 (1)

26-33: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Route the response through the primary webhook client.

endpoint.update() is performed with webhookPrisma, but the returned resource is built from findWebhookEndpointResource, whose ApiWebhookEndpointPresenter reads via webhookReplica. Replica lag can return the previous PAUSED status immediately after enabling; build the response from the updated endpoint row or re-fetch via webookPrisma.

apps/webapp/app/routes/api.v1.webhooks.endpoints.$endpointId.rotate-secret.ts-37-47 (1)

37-47: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Secret rotation isn't fail-safe: a mid-request failure can permanently strand the endpoint.

secretKey (Line 38) is deterministic per endpoint, so setSecret (Line 39) overwrites the endpoint's live signing secret immediately — before the caller has received the response at Line 47. Since prisma (secret store) and webhookPrisma (endpoint row) are separate databases, these writes cannot be made atomic with a single $transaction. If the process fails between line 39 and line 47 (network drop, the webhookPrisma.webhookEndpoint.update throwing, etc.), the new secret is already live, the caller never received it, and there is no way to retrieve it again — the endpoint is effectively bricked until another rotation, which carries the same risk.

🔒 Suggested fix: version the key, swap the pointer, don't overwrite in place
-    const secret = `whsec_${randomBytes(32).toString("hex")}`;
-    const secretKey = `webhook:signing-secret:${endpoint.id}`;
-    await getSecretStore("DATABASE", { prismaClient: prisma }).setSecret(secretKey, { secret });
-    await webhookPrisma.webhookEndpoint.update({
-      where: { id: endpoint.id },
-      data: { signingSecretKey: secretKey },
-    });
+    const secret = `whsec_${randomBytes(32).toString("hex")}`;
+    const secretKey = `webhook:signing-secret:${endpoint.id}:${randomBytes(8).toString("hex")}`;
+    await getSecretStore("DATABASE", { prismaClient: prisma }).setSecret(secretKey, { secret });
+    // Only swap the pointer after the new secret is durably stored; the old secret
+    // (and old key) remain valid verification material until this succeeds.
+    await webhookPrisma.webhookEndpoint.update({
+      where: { id: endpoint.id },
+      data: { signingSecretKey: secretKey },
+    });
+    // TODO: schedule deletion of the previous secretKey after a grace period.
apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks._index/route.tsx-116-133 (1)

116-133: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Presenter failures are silently swallowed as "no deliveries."

.catch(() => ({ deliveries: [], pagination: {} })) masks any ClickHouse/DB failure as an empty result with no logging. Users see "No deliveries for this webhook yet" during an actual outage, and there's no signal for on-call to notice the failure.

🩹 Suggested fix
   const list = await presenter
     .call({...})
-    .catch(() => ({ deliveries: [], pagination: {} }));
+    .catch((error) => {
+      logger.error("Failed to load webhook deliveries", { error });
+      return { deliveries: [], pagination: {}, error: true };
+    });

Then surface error in the UI as a distinct state from "no data yet."

internal-packages/webhook-engine/src/engine/filter/parse.ts-173-180 (1)

173-180: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Field-to-field (valueRef) clauses silently always fail for non-comparison operators.

parseOp accepts word operators (in, startsWith, endsWith, contains, and nin via not in), and the branch below then builds a valueRef clause for any of them when the RHS is a namespace-prefixed word. But applyOpRef in evaluate.ts only implements eq/neq/gt/lt/gte/lte and returns false in its default case — so a filter like event.tags in event.allowed or event.title startsWith event.repository.name parses cleanly yet evaluates to false for every delivery. For a routing/loop-guard filter this is a silent no-match that drops all events.

The comment in evaluate.ts ("Non-comparison ops are rejected by the parser/types") documents the intended contract, but nothing enforces it here. Reject non-comparison ops when building a valueRef clause so authors get a FilterParseError instead of a silently-dead filter.

🐛 Proposed guard
     const operand = peek();
     if (operand?.t === "word" && NAMESPACES.has(operand.v.split(".")[0])) {
       next();
+      if (op !== "eq" && op !== "neq" && op !== "gt" && op !== "lt" && op !== "gte" && op !== "lte") {
+        throw new FilterParseError(
+          `field-to-field comparison only supports ==, !=, >, <, >=, <= (got "${op}")`
+        );
+      }
       return { kind: "clause", path, op, valueRef: operand.v };
     }
internal-packages/webhook-engine/src/engine/verification/hmac.ts-33-37 (1)

33-37: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Reject parsing failures before returning ok: true. parseEventBody() and tryParseJson() return { error } for non-parseable bodies, but hmac.ts and sharedSecret.ts spread that result into the success object, producing ok: true with an error field and no parsedEvent. Check .error and return the verifier’s fail(...) path before building the ok: true result; for sharedSecret.ts, this affects the success path on placement: "header"/bearer/basic secrets, and body secrets already fail to extract when parsing fails.

internal-packages/webhook-engine/src/engine/verification/index.ts-21-41 (1)

21-41: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

verify()'s call graph can throw synchronously, breaking the fail-closed contract relied on by the caller. The engine/index.ts ingest() flow explicitly safeParses the artifact and calls verify() with no try/catch, intending that bad input always becomes { outcome: "verification_failed" } (a 400) rather than a 5xx. Two spots inside the verify() call graph currently violate that assumption by throwing raw Errors instead of returning { ok: false }.

  • internal-packages/webhook-engine/src/engine/verification/index.ts#L21-L41: return { ok: false, error } instead of throwing for the "bundle" kind and the unregistered-scheme default branch.
  • internal-packages/webhook-engine/src/engine/verification/urlSecret.ts#L9-L17: wrap new URL(input.url) in a try/catch and return { ok: false, error: "invalid url", ... } on failure instead of letting it throw.
🟡 Minor comments (24)
docs/webhooks/channels.mdx-11-19 (1)

11-19: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Make the examples self-contained.

  • docs/webhooks/channels.mdx#L11-L19: import streamText from ai and anthropic from @ai-sdk/anthropic.
  • docs/webhooks/human-in-the-loop.mdx#L101-L105: import webhooks from @trigger.dev/sdk.

As per coding guidelines, “Code examples must be complete and runnable where possible.”

Source: Coding guidelines

docs/webhooks/connect.mdx-3-3 (1)

3-3: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Describe this as a verification credential, not always a signing secret.

webhooks.discord() uses the provider’s public key rather than a shared secret, so these universal instructions lead Discord users to configure the wrong value. Distinguish shared-secret and public-key flows.

Also applies to: 15-22

internal-packages/webhook-sources/catalog/mark.ts-33-45 (1)

33-45: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Keep preset, tier, and checklist state consistent.

--preset accepts arbitrary values and does not recompute the tier; --tier sample-only can also retain a preset. Derive tier/checklist from a validated preset, or reject contradictory flag combinations.

internal-packages/webhook-sources/catalog/build-brief.md-24-24 (1)

24-24: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Add a language tag to the TypeScript fence.

This triggers MD040.

Proposed fix
-```
+```typescript

Source: Linters/SAST tools

internal-packages/webhook-sources/catalog/build-v1.ts-33-33 (1)

33-33: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Keep Jira out of the github preset.

Jira Cloud uses X-Hub-Signature with sha256=, while the catalog’s github preset models X-Hub-Signature-256; this makes the generated tier/checklist incorrectly first-class and can break round-trip verification. Set this row’s preset to null until a Jira-specific verifier is added.

internal-packages/webhook-sources/catalog/status.ts-7-18 (1)

7-18: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

sampleCount is never populated, so the CLI always reports "no samples yet".

Provider.sampleCount is read at Line 83, but no entry in providers.json sets this field, so p.sampleCount is always undefined and ${samples} always resolves to "no samples yet" regardless of actual sample coverage — the reported status is misleading for every provider.

🐛 Proposed fix: derive sampleCount instead of trusting an unset JSON field
-const line = (p: Provider) => {
+const line = (p: Provider, sampleCounts: Record<string, number>) => {
   const { done, total } = dodProgress(p);
   const owner = p.owner ? ` owner=${p.owner}` : "";
   const drift = p.status === "complete" && !derivedComplete(p) ? "  [!] checklist incomplete" : "";
-  const samples = p.sampleCount > 0 ? `${p.sampleCount} samples` : "no samples yet";
+  const count = sampleCounts[p.id] ?? 0;
+  const samples = count > 0 ? `${count} samples` : "no samples yet";
Do you want me to wire in an actual sample count (e.g. by importing/counting `src/samples.ts` entries per provider)?

Also applies to: 79-87

internal-packages/webhook-sources/src/registry/postmark.ts-3-17 (1)

3-17: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Point docsUrl at Postmark’s general webhook page.

This registry entry models Postmark delivery-related webhooks with RecordType, but the docsUrl sends users to the separate inbound-webhook documentation. Use the webhook overview URL so the docs match the configured event type.

Suggested fix
-  docsUrl: "https://postmarkapp.com/developer/webhooks/inbound-webhook",
+  docsUrl: "https://postmarkapp.com/developer/webhooks/webhooks-overview",
internal-packages/webhook-sources/src/handAuthored/vapi.ts-20-29 (1)

20-29: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Align epoch timestamps with the ISO timestamps.

The timestamp and artifact message times resolve to July 8, 2026, while the same call’s createdAt and updatedAt are July 14, 2026. Regenerate the epoch-millisecond values from the July 14 timestamps so the sample is internally coherent.

Also applies to: 113-121, 127-147

internal-packages/webhook-sources/src/handAuthored/brex.ts-7-10 (1)

7-10: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Fix the extrapolated USER_UPDATED sample shape.

USER_UPDATED should include the required Brex fields: event_type, user_id, company_id, and updated_attributes. The current sample hides required payload structure and can mislead users building filters or payloads.

internal-packages/webhook-sources/src/handAuthored/openai.ts-19-19 (1)

19-19: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Use a distinct OpenAI event ID per sample.

The batch.completed, response.completed, and realtime.call.incoming samples all reuse evt_685343a1381c819085d44c354e1b330e. OpenAI event objects use id as the unique event identifier, so consumers using this as their dedupe key can treat distinct samples as the same event.

internal-packages/webhook-sources/src/handAuthored/elevenlabs.ts-18-91 (1)

18-91: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Include the audio availability fields in the transcription samples.

Both post_call_transcription examples omit has_audio, has_user_audio, and Has_response_audio, which are part of the ElevenLabs post-call webhook data object. Add realistic boolean values so the catalog samples match current payload shape.

Also applies to: 103-184

internal-packages/webhook-sources/src/handAuthored/linear.ts-19-21 (1)

19-21: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Update the stale webhookTimestamp values.

The millisecond timestamps resolve to July 2025 while the adjacent createdAt values are July 2026. This gives consumers contradictory event times.

Proposed fix
-      webhookTimestamp: 1751380338084,
+      webhookTimestamp: 1782916338084,
...
-      webhookTimestamp: 1751388322391,
+      webhookTimestamp: 1782924322391,
...
-      webhookTimestamp: 1751389329514,
+      webhookTimestamp: 1782925329514,
...
-      webhookTimestamp: 1751361344201,
+      webhookTimestamp: 1782897344201,
...
-      webhookTimestamp: 1751389533802,
+      webhookTimestamp: 1782925533802,

Also applies to: 63-65, 121-124, 138-140, 173-175

internal-packages/webhook-sources/src/roundtrip.test.ts-61-64 (1)

61-64: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Fail when a selected sample has no verifier config.

A sample with presetId: "custom" or a stale provider mapping enters verifiableSamples, then passes without signing or verification because Line 64 returns. Throw instead so the catalog cannot silently lose round-trip coverage.

Proposed fix
       const config = configForSample(sample);
-      if (!config) return;
+      if (!config) {
+        throw new Error(`No verifier config for ${sample.provider} / ${sample.eventType}`);
+      }
internal-packages/webhook-sources/src/handAuthored/close-crm.ts-222-238 (1)

222-238: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Keep the event and payload opportunity IDs consistent.

Line 223 and Line 238 describe different opportunities, so filters based on event.object_id disagree with task code reading event.data.id.

Proposed fix
-          id: "oppo_8H4sjNso7FyBFaeR3RXi5PMJbilfo0c6UPCxsJtEhCO",
+          id: "oppo_7H4sjNso7FyBFaeR3RXi5PMJbilfo0c6UPCxsJtEhCO",
internal-packages/webhook-sources/src/registry/index.ts-65-67 (1)

65-67: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Guard the registry lookup to own keys only.

registry is a normal object, so getProvider("toString") or getProvider("__proto__") returns an inherited value despite the ProviderRegistryEntry | undefined return type. Update the generator to use an own-key check or a null-prototype map.

apps/webapp/app/services/webhookDeliveriesReplicationInstance.server.ts-15-16 (1)

15-16: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Use env.DATABASE_URL instead of reading process.env directly.

DATABASE_URL is already exposed by app/env.server.ts, and this function already uses env for every other variable. Replace the raw const { DATABASE_URL } = process.env read with env.DATABASE_URL.

Source: Path instructions

apps/webapp/app/presenters/v3/WebhookDeliveryDetailPresenter.server.ts-87-90 (1)

87-90: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Use findFirst instead of findUnique.

-        this.replica.sessionRun.findUnique({
+        this.replica.sessionRun.findFirst({
           where: { runId: delivery.runId },
           select: { session: { select: { friendlyId: true, externalId: true } } },
         }),

As per path instructions: "Always use Prisma findFirst instead of findUnique."

Source: Path instructions

apps/webapp/app/presenters/v3/ApiWebhookDeliveryPresenter.server.ts-62-67 (1)

62-67: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Error message omits a valid status value.

The "Allowed" list in the validation error hardcodes pending, processing, succeeded, failed but omits filtered, even though filtered is a valid key in API_STATUS_TO_DB and will be accepted. This misleads API consumers debugging an invalid filter[status] value.

🐛 Proposed fix
-          message: `Invalid status values: ${invalid.join(
-            ", "
-          )}. Allowed: pending, processing, succeeded, failed.`,
+          message: `Invalid status values: ${invalid.join(
+            ", "
+          )}. Allowed: ${Object.keys(API_STATUS_TO_DB).join(", ")}.`,
apps/webapp/app/components/webhookDeliveries/v1/WebhookDeliveryFilters.tsx-39-44 (1)

39-44: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Status filter is missing the FILTERED option.

deliveryStatuses only covers 4 of the 5 WebhookDeliveryStatus values. The delivery timeline builder (and this PR's own test suite) treats FILTERED as a first-class outcome, and the server-side loader already accepts it via Object.values(WebhookDeliveryStatus) — but this dropdown has no way to select it, so users can't filter deliveries down to filtered-out events through the UI.

 const deliveryStatuses: { value: WebhookDeliveryStatus; title: string; color: string }[] = [
   { value: "PENDING", title: "Pending", color: "`#878C99`" },
   { value: "PROCESSING", title: "Processing", color: "`#3B82F6`" },
   { value: "SUCCEEDED", title: "Succeeded", color: "`#28BF5C`" },
   { value: "FAILED", title: "Failed", color: "`#E11D48`" },
+  { value: "FILTERED", title: "Filtered", color: "`#878C99`" },
 ];

Since the comment says this list intentionally "Match[es] DeliveriesTable's DELIVERY_STATUS_COLOR / DELIVERY_STATUS_LABEL," worth checking whether that file (not in this batch) also needs the same addition.

apps/webapp/app/components/webhookDeliveries/v1/DeliveriesTable.tsx-240-247 (1)

240-247: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

"View session" popover item uses the run color, not the session color.

Elsewhere in this file (Line 160-161) sessions use text-sessions; this menu item uses text-runs for "View session", inconsistent with the established color-coding convention distinguishing sessions from runs.

🎨 Fix
           {sessionPath ? (
             <PopoverMenuItem
               to={sessionPath}
               icon={ArrowRightIcon}
-              leadingIconClassName="text-runs"
+              leadingIconClassName="text-sessions"
               title="View session"
             />
           ) : null}
apps/webapp/app/components/webhookConsole/SampleSourcePicker.tsx-56-62 (1)

56-62: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Category search is case-sensitive; label search isn't.

query is lowercased but p.category is compared as-is, so a category like "Payments" won't match a lowercase search term even though the label check does lowercase.

🐛 Fix
-      (p) => p.label.toLowerCase().includes(query) || (p.category ?? "").includes(query)
+      (p) => p.label.toLowerCase().includes(query) || (p.category ?? "").toLowerCase().includes(query)
packages/core/src/v3/schemas/schemas.ts-2-8 (1)

2-8: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Restore type-only imports for RequireKeys, AnyRunTypes, and the inferred RetrieveRunResponse.

These symbols are export type utilities/inferences, not runtime exports; importing them without type can break with isolatedModules/verbatimModuleSyntax-style builds that cannot drop unused imports.

  • packages/core/src/v3/schemas/schemas.ts#L2-L8
  • packages/core/src/v3/types/index.ts#L1-L3
packages/core/src/v3/schemas/webhookApi.ts-64-70 (1)

64-70: 🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

pagination is required here but documented as optional (Line 7). A deliveries response serialized as { data } (no pagination key) would fail ListWebhookDeliveriesResponse.parse(...). Consider making it optional to match the documented { data, pagination? } contract and avoid a parse failure.

🛡️ Proposed change
 export const ListWebhookDeliveriesResponse = z.object({
   data: z.array(WebhookDeliveryListItem),
   pagination: z.object({
     next: z.string().optional(),
     previous: z.string().optional(),
-  }),
+  }).optional(),
 });
internal-packages/webhook-engine/src/engine/verification/urlSecret.ts-9-17 (1)

9-17: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Unguarded new URL(input.url) can throw on malformed input.

If input.url isn't a valid absolute URL, new URL() throws synchronously and isn't caught here or (per the cross-file engine/index.ts ingest() evidence) by the caller, risking an unhandled 5xx instead of a fail-closed { ok: false } result — the same fail-closed contract violated in verification/index.ts.

🛡️ Proposed fix
   verify(config, input): VerifierResult {
     const cfg = config as Extract<UrlSecretConfig, { scheme: "url-secret" }>;
-    const u = new URL(input.url);
+    let u: URL;
+    try {
+      u = new URL(input.url);
+    } catch {
+      return { ok: false, error: "invalid url", idempotencyKey: derive0(cfg, input) };
+    }
🧹 Nitpick comments (13)
internal-packages/replication/src/client.ts (1)

34-39: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use a type alias for this configuration shape.

LogicalReplicationClientOptions is a data shape, so extend a type alias rather than an interface.

Proposed change
-export interface LogicalReplicationClientOptions {
+export type LogicalReplicationClientOptions = {
   // ...
-}
+};

Source: Coding guidelines

internal-packages/webhook-sources/catalog/providers.json (1)

23-923: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Catalog omits 11 registry providers that exist in code.

registry/adyen.ts, bigcommerce.ts, bitbucket.ts, checkout.ts, commercelayer.ts, monday.ts, paddlebilling.ts, paddleclassic.ts, paypal.ts, pipedrive.ts, and woocommerce.ts all exist per the cohort's file list, but none of them have a corresponding entry in this providers array. catalog/status.ts derives "remaining"/"claimable"/"complete" counts purely from this file, so these 11 providers are silently excluded from progress tracking (the tool will report the wave as fully done while 11 shipped providers are untracked).

internal-packages/webhook-sources/src/handAuthored/slack.ts (1)

17-17: 🔒 Security & Privacy | 🔵 Trivial | 💤 Low value

Replace the credential-shaped Slack token with an explicit fixture value.

The shared-secret callback token is repeated in several payloads; use an unmistakable placeholder such as "fixture-token" across these samples.

Sources: MCP tools, Linters/SAST tools

apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.test.tasks.$taskParam/route.tsx (1)

134-138: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use v3WebhookTaskPath instead of a hand-built URL.

This PR adds v3WebhookTaskPath in pathBuilder.ts (which also encodeURIComponents the slug); this redirect duplicates that logic with raw string interpolation instead. See consolidated comment.

apps/webapp/app/routes/resources.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks.endpoints.$endpointParam.send.ts (1)

175-176: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use v3WebhookDeliveryPath instead of a hand-built URL.

Duplicates the new v3WebhookDeliveryPath pathBuilder helper with raw string interpolation. See consolidated comment.

apps/webapp/app/components/webhookConsole/WebhookComposer.tsx (1)

117-119: 🎯 Functional Correctness | 🔵 Trivial | 💤 Low value

signatureMode is not reconciled when the selected endpoint changes.

signatureMode is initialized once from signedAvailable. When the user switches endpoints (dropdown at Lines 267-287) to one that is asymmetric or has no signing secret, a previously-selected "signed" mode stays set even though its SelectItem becomes disabled, so a subsequent send submits signed for an unsupported endpoint. Consider resetting to "simulate" in an effect when signedAvailable becomes false.

apps/webapp/app/v3/services/createBackgroundWorker.server.ts (1)

225-232: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Non-ServiceValidationError webhook-sync failures are swallowed (deploy reports success).

Unlike syncDeclarativeSchedules above (which wraps unexpected errors into a ServiceValidationError and rethrows), a transient failure here (e.g. a webhookPrisma/prisma error) is only logged; the deploy then completes as successful with webhook endpoints left unsynced/partially updated. Consider mirroring the schedules path so an infra failure fails the deploy rather than silently diverging routing state.

apps/webapp/app/presenters/v3/WebhookDetailPresenter.server.ts (1)

447-450: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Prefer the structured logger over console.error (also at Line 541).

console.error bypasses Logger.onError forwarding (e.g. Sentry) and structured fields used elsewhere in the presenters. Consider importing logger from ~/services/logger.server for these ClickHouse query-failure paths.

apps/webapp/app/presenters/v3/ApiWebhookDeliveryPresenter.server.ts (1)

32-42: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate delivery-mapping logic.

ApiWebhookDeliveryPresenter.call re-implements the same field mapping already encapsulated in toApiListItem, risking silent drift if one is updated without the other.

♻️ Proposed refactor
       return {
-        id: d.friendlyId,
-        webhook: d.webhook?.slug ?? null,
-        status: DB_STATUS_TO_API[d.status],
-        externalDeliveryId: d.externalDeliveryId,
-        runId: d.run?.friendlyId ?? null,
-        createdAt: d.createdAt,
-        processedAt: d.processedAt,
+        ...toApiListItem(d),
         idempotencyKey: d.idempotencyKey,
         event: d.parsedEvent ?? null,
         headers: (d.headers as Record<string, string> | null) ?? null,
         rawBodyHash: d.rawBodyHash,
         error: d.errorMessage,
         filterReason: d.filterReason,
         updatedAt: d.updatedAt,
       };

Note: toApiListItem takes WebhookDeliveryListItem; d here is a WebhookDeliveryDetail — confirm it's structurally compatible (a superset) before applying.

Also applies to: 133-148

apps/webapp/app/routes/_app.orgs.$organizationSlug.projects.$projectParam.env.$envParam.webhooks.deliveries.$deliveryParam/route.tsx (1)

109-115: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Duplicate duration-formatting logic.

This local formatDuration re-implements what formatDuration from @trigger.dev/core/v3/utils/durations already provides (used in apps/webapp/app/components/webhookDeliveries/v1/DeliveryTimeline.tsx). Consider reusing the shared utility for consistent formatting across the delivery UI.

internal-packages/webhook-engine/src/engine/types.ts (1)

28-58: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Prefer a type alias for the WebhookEngineOptions data shape.

WebhookEngineOptions is a plain configuration/object shape, so the repo convention of types-over-interfaces applies (the two callback interfaces are behavioral contracts and are fine as-is).

♻️ Convert to a type alias
-export interface WebhookEngineOptions {
+export type WebhookEngineOptions = {
   logger?: Logger;
   ...
   deliverToSession?: DeliverWebhookToSessionCallback;
-}
+};

As per coding guidelines: "Use types over interfaces for TypeScript".

Sources: Coding guidelines, Learnings

internal-packages/webhook-engine/src/engine/index.ts (1)

70-93: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Remove the redundant /milliseconds suffix from the duration metric.

The histogram is declared with unit: "ms" and Prometheus-exported metric names already receive the _milliseconds unit suffix, so these records are exported as webhook_delivery_execution_duration_milliseconds_milliseconds. Rename it to webhook_delivery_execution_duration and keep the "ms" unit; update dashboards/queries accordingly.

Source: Coding guidelines

internal-packages/webhook-engine/src/engine/verification/sharedSecret.ts (1)

25-48: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Loosely-typed cfg: any in extractCandidate/failSS.

Unlike hmac.ts's fail(error, cfg: HmacConfig, input), these helpers type cfg as any, losing compile-time checks on cfg.placement, cfg.fieldName, and cfg.idempotencyField in secret-extraction code.

♻️ Proposed fix: use the narrowed config type
-function extractCandidate(cfg: any, input: VerifyInput): string | undefined {
+function extractCandidate(
+  cfg: Extract<SharedSecretConfig, { scheme: "shared-secret" }>,
+  input: VerifyInput
+): string | undefined {
-function failSS(error: string, cfg: any, input: VerifyInput): VerifierResult {
+function failSS(
+  error: string,
+  cfg: Extract<SharedSecretConfig, { scheme: "shared-secret" }>,
+  input: VerifyInput
+): VerifierResult {

Also applies to: 50-62

@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from 65e81fc to 37d5285 Compare July 23, 2026 09:32
@pkg-pr-new

pkg-pr-new Bot commented Jul 23, 2026

Copy link
Copy Markdown

Open in StackBlitz

@trigger.dev/build

npm i https://pkg.pr.new/@trigger.dev/build@7a609bf

trigger.dev

npm i https://pkg.pr.new/trigger.dev@7a609bf

@trigger.dev/core

npm i https://pkg.pr.new/@trigger.dev/core@7a609bf

@trigger.dev/python

npm i https://pkg.pr.new/@trigger.dev/python@7a609bf

@trigger.dev/react-hooks

npm i https://pkg.pr.new/@trigger.dev/react-hooks@7a609bf

@trigger.dev/redis-worker

npm i https://pkg.pr.new/@trigger.dev/redis-worker@7a609bf

@trigger.dev/rsc

npm i https://pkg.pr.new/@trigger.dev/rsc@7a609bf

@trigger.dev/schema-to-json

npm i https://pkg.pr.new/@trigger.dev/schema-to-json@7a609bf

@trigger.dev/sdk

npm i https://pkg.pr.new/@trigger.dev/sdk@7a609bf

commit: 7a609bf

@ericallam ericallam changed the title feat: hosted webhook ingress and chat.agent channels feat: hosted webhooks, agent channels, and human-in-the-loop Jul 23, 2026
@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch 2 times, most recently from 0375d33 to 8c3eeb6 Compare July 27, 2026 16:20
samejr added a commit that referenced this pull request Aug 3, 2026
A UI pass over the webhooks dashboard, on top of #4344. No behaviour
changes beyond the fixes below.

## Deliveries list

- Whole row is clickable. The external delivery ID, created, processed
and error cells had no link, and the target cell only linked when the
delivery had a run or session, so most of each row was dead.
- Dimmed "None" and "Unknown" cells now brighten with the row on hover.
- The new-deliveries button sits inline, left of the pager, instead of
on its own row beneath it.
- The table scrolls. It was passing `stickyHeader`, which switches the
table container to `overflow-visible` and stops it being the scroll
container; every other list in the app leaves it off. The header stays
sticky either way.
- 60 deliveries per page, up from 25. Test tag uses the shared `Badge`,
the webhook icon matches the Tasks page, and the Status and More filters
menus drop their redundant search fields.

## Delivery detail

- Dropped the duplicate status badge from the title bar; the sidebar
already has a Status row.
- The "nothing was captured" tab messages are centred and a size larger.
- Copyable sidebar values ellipsise instead of overflowing their column,
so an unbreakable hash or opaque id no longer runs past the edge.
`CopyableText` gains an opt-in `truncate` prop that reserves a gutter
for the copy button.
- The delivery timeline's thick bar is rounded at the top. The run
timeline gets that corner from the `start-cap-thick` event above its
thick line, but a delivery only has two timestamps, so the line itself
starts the bar and had a square top on every succeeded and failed
delivery. `RunTimelineLine` gains an opt-in `roundedTop`, so other
callers are unaffected.

## Navigation

Webhooks was a section containing a single item. It now sits as a
top-level item below Sessions, and the page is titled "Webhook
deliveries". Registering the page in the favourites registry also fixes
its favourite name, which was saving as "Page: Deliveries".

## Also

One fix outside the UI: the delivery seed script minted `id` and
`friendlyId` as two independent ids, but the detail lookup derives the
row id from the friendlyId, so every seeded delivery's page reported
that it could not be found.
ericallam pushed a commit that referenced this pull request Aug 3, 2026
A UI pass over the webhooks dashboard, on top of #4344. No behaviour
changes beyond the fixes below.

## Deliveries list

- Whole row is clickable. The external delivery ID, created, processed
and error cells had no link, and the target cell only linked when the
delivery had a run or session, so most of each row was dead.
- Dimmed "None" and "Unknown" cells now brighten with the row on hover.
- The new-deliveries button sits inline, left of the pager, instead of
on its own row beneath it.
- The table scrolls. It was passing `stickyHeader`, which switches the
table container to `overflow-visible` and stops it being the scroll
container; every other list in the app leaves it off. The header stays
sticky either way.
- 60 deliveries per page, up from 25. Test tag uses the shared `Badge`,
the webhook icon matches the Tasks page, and the Status and More filters
menus drop their redundant search fields.

## Delivery detail

- Dropped the duplicate status badge from the title bar; the sidebar
already has a Status row.
- The "nothing was captured" tab messages are centred and a size larger.
- Copyable sidebar values ellipsise instead of overflowing their column,
so an unbreakable hash or opaque id no longer runs past the edge.
`CopyableText` gains an opt-in `truncate` prop that reserves a gutter
for the copy button.
- The delivery timeline's thick bar is rounded at the top. The run
timeline gets that corner from the `start-cap-thick` event above its
thick line, but a delivery only has two timestamps, so the line itself
starts the bar and had a square top on every succeeded and failed
delivery. `RunTimelineLine` gains an opt-in `roundedTop`, so other
callers are unaffected.

## Navigation

Webhooks was a section containing a single item. It now sits as a
top-level item below Sessions, and the page is titled "Webhook
deliveries". Registering the page in the favourites registry also fixes
its favourite name, which was saving as "Page: Deliveries".

## Also

One fix outside the UI: the delivery seed script minted `id` and
`friendlyId` as two independent ids, but the detail lookup derives the
row id from the friendlyId, so every seeded delivery's page reported
that it could not be found.
@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from a0dcd78 to f6b224b Compare August 3, 2026 15:54
@ericallam
ericallam marked this pull request as ready for review August 3, 2026 15:54
devin-ai-integration[bot]

This comment was marked as resolved.

ericallam pushed a commit that referenced this pull request Aug 7, 2026
A UI pass over the webhooks dashboard, on top of #4344. No behaviour
changes beyond the fixes below.

## Deliveries list

- Whole row is clickable. The external delivery ID, created, processed
and error cells had no link, and the target cell only linked when the
delivery had a run or session, so most of each row was dead.
- Dimmed "None" and "Unknown" cells now brighten with the row on hover.
- The new-deliveries button sits inline, left of the pager, instead of
on its own row beneath it.
- The table scrolls. It was passing `stickyHeader`, which switches the
table container to `overflow-visible` and stops it being the scroll
container; every other list in the app leaves it off. The header stays
sticky either way.
- 60 deliveries per page, up from 25. Test tag uses the shared `Badge`,
the webhook icon matches the Tasks page, and the Status and More filters
menus drop their redundant search fields.

## Delivery detail

- Dropped the duplicate status badge from the title bar; the sidebar
already has a Status row.
- The "nothing was captured" tab messages are centred and a size larger.
- Copyable sidebar values ellipsise instead of overflowing their column,
so an unbreakable hash or opaque id no longer runs past the edge.
`CopyableText` gains an opt-in `truncate` prop that reserves a gutter
for the copy button.
- The delivery timeline's thick bar is rounded at the top. The run
timeline gets that corner from the `start-cap-thick` event above its
thick line, but a delivery only has two timestamps, so the line itself
starts the bar and had a square top on every succeeded and failed
delivery. `RunTimelineLine` gains an opt-in `roundedTop`, so other
callers are unaffected.

## Navigation

Webhooks was a section containing a single item. It now sits as a
top-level item below Sessions, and the page is titled "Webhook
deliveries". Registering the page in the favourites registry also fixes
its favourite name, which was saving as "Page: Deliveries".

## Also

One fix outside the UI: the delivery seed script minted `id` and
`friendlyId` as two independent ids, but the detail lookup derives the
row id from the friendlyId, so every seeded delivery's page reported
that it could not be found.
@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from f6b224b to baf5e98 Compare August 7, 2026 22:53
@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Observability map

As of ead2f98.

18/100 over 431 measured of 447 entry points (base 18, no change)

What this PR changed

route base head now failing
/_app/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/_index new 0 request-context
/_app/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/:webhookParam new 0 request-context
/_app/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/deliveries/:deliveryParam new 0 request-context
/_app/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/endpoints/:endpointParam new 0 request-context
/api/v1/projects/:projectRef/init new 0 error-classification, request-context
/api/v1/webhooks/deliveries new 0 request-context
/api/v1/webhooks/deliveries/:deliveryId new 0 request-context
/api/v1/webhooks/deliveries/:deliveryId/replay new 0 request-context
/api/v1/webhooks/endpoints new 0 request-context
/api/v1/webhooks/endpoints/:endpointId new 0 request-context
/api/v1/webhooks/endpoints/:endpointId/disable new 0 request-context
/api/v1/webhooks/endpoints/:endpointId/enable new 0 request-context
/api/v1/webhooks/endpoints/:endpointId/rotate-secret new 0 request-context
/resources/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/deliveries/live new 0 request-context
/resources/orgs/:organizationSlug/projects/:projectParam/env/:envParam/webhooks/endpoints/:endpointParam/replay-source new 0 request-context

and 3 more

FIX FIRST

  • /api/v1/projects/:projectRef/envvars (sensitive) - auth-boundary, request-context
  • /auth/sso (sensitive) - auth-boundary, request-context
  • /_app/orgs/:organizationSlug/settings/team (sensitive) - error-classification, auth-scope, request-context

AUDIT 3 of 50 sensitive mutations record an actor. 47 without one.
CONTEXT 11 of 431 entry points name a tenant on a failure path. 342 appear only here, 39 of them sensitive, in the JSON rather than the fix list.

What the score is made of
CHECKS
  error-classification  169 applicable,  94 pass,   0 sole, global without it 9
  auth-boundary          62 applicable,  57 pass,   0 sole, global without it 14
  auth-scope             19 applicable,  17 pass,   0 sole, global without it 17
  request-context       431 applicable,  11 pass, 240 sole, global without it 63
  audit-trail            50 applicable,   3 pass,   0 sole, not in the score

The score and findings here are report-only and never gate the merge. Separately, a required test suite keeps this tool's symbol and route lists in sync with the code they name, and can fail a pull request that renames or removes a symbol they reference, or that adds the first route with a segment they anticipate. Each failure names the list to edit. The rules and their reasons: internal-packages/observability-map/README.md.

devin-ai-integration[bot]

This comment was marked as resolved.

@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from baf5e98 to d2004d2 Compare August 8, 2026 06:25
@ericallam ericallam changed the title feat: hosted webhooks, agent channels, and human-in-the-loop feat(webapp): hosted webhook ingress, delivery pipeline, and dashboard Aug 8, 2026
@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from d2004d2 to de4fce8 Compare August 8, 2026 08:20
devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from 8ee8900 to 648409d Compare August 9, 2026 19:10
Webhook events larger than 8KB reached the task as a { truncated, bytes } placeholder instead of the real payload: the stored event was capped at 8KB, and that same column is routed as the run payload for task, session, and replay deliveries.

The verified event is now stored and routed in full (bounded by the ingress body-size limit). The delivery detail view caps the payload it renders for readability and notes that the full event was delivered to the task.
A redeploy no longer re-activates a hosted webhook endpoint that was disabled via the API: the declarative sync only marks an endpoint active when it first creates it, so an operator disable survives future deploys. A deploy that omits the webhook list entirely (an older client) also no longer deactivates existing endpoints, which is now distinguished from an explicit empty list.
…secret change

Rotating or generating a webhook signing secret from the dashboard now invalidates the engine's cached endpoint immediately, so deliveries are verified against the new secret right away instead of being rejected for up to the cache TTL. The HTTP API route already did this; the dashboard generate and set/rotate actions now match.
devin-ai-integration[bot]

This comment was marked as resolved.

devin-ai-integration[bot]

This comment was marked as resolved.

@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from 840bf41 to ed87083 Compare August 11, 2026 09:31
…il lookup

The delivery detail point lookup queried WebhookDelivery by id and environment
only. The table is RANGE-partitioned on createdAt, so with no createdAt
predicate Postgres cannot prune and probes every daily partition.

The delivery id now embeds its mint timestamp (a 6-byte big-endian unix ms
prefix plus random bytes, base32hex encoded), and the engine stores that same
timestamp as the row's createdAt, so the id's timestamp is the partition key.
getDelivery recovers it from the friendlyId and adds it as an exact predicate,
pruning to the row's partition for every caller.
… disables

A webhook removed from the deploy manifest is deactivated by the declarative
sync. On a later deploy that re-declares it, the endpoint stayed INACTIVE and
silently dropped deliveries. A new WebhookEndpoint.manuallyDeactivatedAt
timestamp distinguishes an operator disable from that auto-deactivation: the
sync reactivates an auto-deactivated endpoint on re-declare, but leaves an
operator-disabled one alone. The disable and enable endpoints set and clear the
timestamp.
…e id and pager refresh

Three dashboard fixes:

The test-send action now returns early when WEBHOOK_ENABLED is off, matching the
ingress route, so a test send cannot record a delivery the disabled engine would
never process (no partition, no worker).

The duplicate outcome no longer re-prefixes the delivery id (it is already a
friendlyId), so the console shows a valid id and a working "view original" link
instead of a whd_whd_ id.

The "new deliveries" button now clears the deliveriesCursor/deliveriesDirection
params this page actually paginates on, so it shows the new rows past page one.
WEBHOOK_WORKER_ENABLED defaults from WORKER_ENABLED and follows the same
true/false convention as the other worker switches, but the engine compared it
to "0" (the convention of a different set of flags). A web-only instance
(WORKER_ENABLED=false) therefore kept running webhook delivery jobs once the
feature was enabled. Compare against "true" like its peers.
An earlier change on this branch inadvertently reverted the dev postgres setup
back to plain postgres:14, dropping the pg_partman build and its
shared_preload_libraries that were added separately on main. Restore
docker-compose.yml, dev-compose.yml, and Dockerfile.postgres to match main so
merging this branch does not undo that.
…otency per endpoint

Two delivery correctness fixes.

The ensurePartitions cron is only enqueued for its next scheduled tick and the
table has no default partition, so a freshly enabled engine had no partition for
incoming events until the first nightly run and rejected them. Run the same
ensurePartitions pass once at startup.

The Run Engine idempotency key was the raw provider delivery id, but Run Engine
idempotency is scoped per task and environment while the front gate is scoped
per endpoint. Two endpoints in one environment routing the same event id to the
same task could collapse into one run. Prefix the key with the endpoint id to
match the front gate.
…ILTERED deliveries

The delivery detail presenter looked up the backing session with findUnique,
which the webapp avoids for its query-batching defects; switch to findFirst. The
per-webhook activity chart's status series omitted FILTERED even though it is a
first-class delivery status with a color everywhere else, so filtered deliveries
vanished from the chart; add it to the series.
The "N new deliveries" badge queried the live count with only the endpoint and
time window, dropping the status, webhook, delivery-id and test filters the list
was showing, so on a filtered list it announced deliveries the list would never
display. The count now runs through the same filter resolution as the list
(WebhookDeliveriesListPresenter.countNewDeliveries), and the live-reload hook
forwards the active filter params, so the badge only counts rows the list shows.
Running ensurePartitions at engine startup fired it on every worker instance at
boot, and createPartition (partitionExists then CREATE, with no lock) is not
safe to run concurrently across instances. The nightly cron already maintains
partitions from a single consumer, which is safe. Initial bootstrap will move to
an admin-triggered action rather than running on every instance at startup.
…s window to 7 days

The delivery detail page called useState after its not-found early return, so
navigating between a missing and an existing delivery changed the hook count and
crashed the page. The hook now runs unconditionally above the return.

The per-webhook and per-endpoint deliveries lists passed period through unset, so
they listed every delivery ever while the time control showed "Last 7 days" (and
scanned all history). They now default to 7 days when no explicit window is set,
matching the top-level list.
…ed duplicate

On the front-gate duplicate branch, when the stored gate value had already
expired between the failed set-NX and the get, the code returned this request's
freshly generated friendlyId, which points to a delivery that was never created
(a 404 when opened). The duplicate outcome's deliveryId is now optional and
returns the stored id or nothing; the console skips the link/redirect when it is
absent.
TimeFilter clears the generic cursor/direction on apply, but the webhook detail
page paginates under deliveriesCursor/deliveriesDirection and runsCursor/
runsDirection, so changing the range left a stale page and the list came back
empty or misaligned. TimeFilter gains an optional clearParams, and the page
passes its namespaced pagination params so a range change returns to the first
page.
…days

getDeliveriesByFriendlyIds looked rows up by id only, so the live poll probed
every retained daily partition every few seconds per open dashboard. The ids are
time-encoded, so when they all decode we bound the query to the span of their
mint timestamps (which equal createdAt), pruning to the visible page's few days.
A legacy id in the set falls back to the unbounded lookup.
…nvalidation

invalidateEndpoint only cleared the in-process cache, so a secret rotation or
enable/disable took effect at once on the handling instance while every other
instance still served the stale entry until the cache TTL expired. That uneven
convergence is more confusing than useful, so remove it: all instances now
converge uniformly within WEBHOOK_ENDPOINT_CACHE_TTL_MS. A cross-instance
invalidation channel can come later if a shorter window is needed.
…ds calc

getDeliveriesByFriendlyIds computed the createdAt prune window in three passes
(map to timestamps, filter, then map again inside Math.min(...) / Math.max(...)).
Extract it to a single-pass deliveryIdsCreatedAtBounds that decodes each id once
and tracks min/max in one loop, with no Math.min(...spread), which builds the
whole argument list and overflows the stack on large inputs. Same result, plus a
unit test covering the span, single-id, empty, and legacy-fallback cases.
The delivery-id bounds calc decoded every id in the set to find the createdAt
span. Delivery id bodies are base32hex(big-endian timestamp then random bytes),
and base32hex is order-preserving, so lexical order equals chronological order:
the earliest and latest timestamps sit at the lexical extremes. Track the min
and max id in a single pass and decode only those two.

The list-hydration query shares the same single-pass min/max helper for its
already-decoded page timestamps, dropping the last Math.min(...spread) in the
delivery repository (the spread overflows the call stack on large inputs).
@ericallam
ericallam force-pushed the feat/hosted-webhook-ingress branch from ed87083 to 7a609bf Compare August 11, 2026 09:38
devin-ai-integration[bot]

This comment was marked as resolved.

Two configs were accepted but would then fail-close every delivery.

The dashboard secret-generation action minted a shared secret for any endpoint,
including asymmetric (public-key) ones, overwriting the stored public key. It now
rejects generation for asymmetric endpoints, matching the public API route.

url-secret verification with path placement can never match on the hosted ingress
URL, whose last path segment is the fixed opaque endpoint id, so deploy-sync now
rejects it with a clear error instead of letting every inbound event 400.
…oading

Two webhook dashboard views showed a misleading state.

An empty deliveries list always rendered 'No deliveries match these filters'
because the detail routes default the window to 7 days, so the presenter treated
that default as an active filter. It now derives hasFilters from whether the user
explicitly set a window, so a webhook that has never received anything shows the
real empty state.

The sample-event picker treated the provider list as loaded before its fetch had
started, flashing 'No providers' on first open. It now shows the spinner until
the data arrives.
devin-ai-integration[bot]

This comment was marked as resolved.

…is down

The per-IP ingress limiter is an async Express middleware, and Express 4 does
not catch a rejected promise from a handler. If the rate-limiter backend is
unreachable, ipLimiter.limit() rejects, next() is never called, and the request
hangs until the client times out. Wrap the check in try/catch and let the
request through on a limiter error (the per-endpoint limiter is the real
protection), matching the OTLP ingress limiter.
The left nav reads only org-level feature flags, so a global FeatureFlag row
makes the pages reachable by URL but does not reveal the nav section for a
non-admin. Point the onboarding step at the org-level flag.
devin-ai-integration[bot]

This comment was marked as resolved.

…ure is disabled

The engine singleton is constructed at boot, opening the front-gate Redis client
and the worker's queue (a second Redis client plus queue-size gauges) regardless
of WEBHOOK_ENABLED. Only worker.start() was flag-gated, so a deployment with
webhooks off still held two Redis connections and polled queue size, and one
without the webhook Redis configured would spam reconnect errors.

Add an engine-level disabled option (set from WEBHOOK_ENABLED !== '1') that skips
building the Redis clients and the worker entirely. The public entry points
(ingest, simulateInject, replayDelivery, getJob) assert the engine is enabled and
quit() no-ops, so an off deployment is genuinely inert. worker.disabled is
unchanged: an ingress-only instance still builds the engine but does not start the
worker loop.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants