From 3b03c4b9ebe9349bedb719ab871a671fa71c983f Mon Sep 17 00:00:00 2001
From: Dotta <34892728+cryppadotta@users.noreply.github.com>
Date: Fri, 11 Sep 2026 10:15:00 -0500
Subject: [PATCH] perf(ui): keep long task chat responsive during streaming
(#13229)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
## Thinking Path
> - Paperclip lets operators manage agent work through tasks.
> - Task chat keeps responses and run history together.
> - A live update rendered every historical bubble and hidden tool row
again.
> - New image callbacks also forced unchanged markdown to parse again.
> - Long conversations saturated the browser main thread.
> - This change reuses unchanged history and mounts folded tools on
first inspection.
> - Operators can read and reply while work continues.
## Linked Issues or Issue Description
**What happened?**
Chat-style tasks with substantial scrollback became almost unusable. A
deterministic browser reproduction with 200 long responses and 4,000
tools consumed 97.8% of the main thread during live updates. It
delivered only 8 updates during the sample.
**Expected behavior**
The task should remain responsive during streaming. Historical markdown
and unopened run details should not repeat expensive render work.
**Steps to reproduce**
1. Install dependencies with `pnpm install`.
2. Run `pnpm exec playwright test --config
tests/perf/task-chat/playwright.config.ts`.
3. Compare the attached performance JSON. The new test fails against the
original rendering code.
**Paperclip version or commit**
Reproduced at `a05b828bc`. The branch is rebased on current master.
**Deployment mode**
Local Chromium and Vite with deterministic fixtures. No database or
agent credentials are required.
Related work: #10463 reduces the issue-page bundle. This change
addresses repeated rendering after the page loads. No duplicate
scrollback fix was found.
## What Changed
- Keep the bubble image callback stable so unchanged markdown can skip
parsing.
- Memoize the settled history separately from the header and streaming
tail. Keep the brief renderer and default attachment array stable.
- Mount folded tool history on first expansion. Keep it mounted
afterward to preserve child state and closing motion. Runtime request
receipts remain visible.
- Add deterministic rendering tests to the normal Vitest suite and an
opt-in Chromium regression fixture.
- Document the reproduction, commands, scope, and local measurements.
## Verification
- Browser reproduction: 97.8% main-thread utilization before; 11.6%
after for tail-only updates; 31.5% after when projection recreates
history objects. Both fixed cases delivered 32 updates.
- Browser checks pass for scroll-position retention, typing, return to
latest, tool inspection, and retained expansion state.
- Focused component suite: 155 tests passed. Post-rebase
thread/performance rerun: 101 tests passed.
- Recursive typecheck, build, Storybook build, and token gates passed.
UI typecheck passed after the final edits.
- Local UI/CLI lane: 5,910 tests passed; ten files hit worker-start
timeouts, then all ten passed with two workers (20 tests).
Shared/adapter lane: 3,128 tests passed.
- Full local `pnpm test:run` was attempted and is **not green**: its
server lane recorded 10,517 passes and 11 failures plus fixture/setup
errors. Queue (31 tests), Cursor/Git-load (9 tests), and missing-binary
failures cleared on isolated reruns / building the runner test binaries.
Two native suites still cannot initialize embedded PostgreSQL on this
host.
- The remaining native-session recovery assertion was reproduced in a
clean worktree at base `a20ecce40` (1 failed, 67 passed across the
native/queue suites). It expects a settled-session error but receives a
semantic-tool-input digest error. No server or runner files changed in
this PR.
- All substantive CI jobs have passed, including build, typecheck, all
server/workspace test shards, all three e2e shards, and the canary dry
run. The unchanged Slack ordering test exhausted its one-second wait on
the first run; its shard passed on rerun. Final aggregate verification
passed: **31 passing checks**, no failures or pending checks; two
optional Storybook deployment/visual checks were skipped. Greptile is
**5/5**, with no unresolved review threads.
- The supplemental local serialized route run was stopped after CI
passed all five serialized shards; it had reported no failures.
- The browser fixture uses the real chat rendering components with a
plain textarea. It does not test the full composer or server transport.
Timing results are local samples; deterministic render-count tests
provide normal CI coverage.
## Risks
- Closed tool content becomes available to DOM search only after first
expansion. Visible run summaries and runtime request receipts remain
available immediately.
- Opened run history remains mounted to preserve child state. The first
full markdown render still scales with conversation size.
- Memo dependencies must stay current when adding render inputs. Tests
check content edits and replacement gallery callbacks.
- No API, database, or permission changes.
## Model Used
OpenAI GPT-6 (Codex). Exact serving variant and context-window size are
not exposed in this session. Used reasoning, repository tools, code
execution, and browser testing. No sub-agents were used.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted/UI/workspace
checks pass; the full local server suite has the baseline/host failures
documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip
---
tests/perf/task-chat/README.md | 44 +++++
tests/perf/task-chat/playwright.config.ts | 15 ++
tests/perf/task-chat/scrollback.spec.ts | 56 ++++++
ui/src/components/TaskChatThread.tsx | 11 +-
.../components/task-chat/TaskChatBubble.tsx | 8 +-
.../TaskChatThreadView.performance.test.tsx | 77 ++++++++
.../task-chat/TaskChatThreadView.tsx | 173 ++++++++++--------
.../task-chat/TaskChatTurn.test.tsx | 4 +-
ui/src/components/task-chat/TaskChatTurn.tsx | 6 +-
ui/src/fixtures/TaskChatPerfHarness.tsx | 57 ++++++
ui/tests/task-chat-perf.html | 5 +
11 files changed, 367 insertions(+), 89 deletions(-)
create mode 100644 tests/perf/task-chat/README.md
create mode 100644 tests/perf/task-chat/playwright.config.ts
create mode 100644 tests/perf/task-chat/scrollback.spec.ts
create mode 100644 ui/src/components/task-chat/TaskChatThreadView.performance.test.tsx
create mode 100644 ui/src/fixtures/TaskChatPerfHarness.tsx
create mode 100644 ui/tests/task-chat-perf.html
diff --git a/tests/perf/task-chat/README.md b/tests/perf/task-chat/README.md
new file mode 100644
index 0000000000..b5de1eb999
--- /dev/null
+++ b/tests/perf/task-chat/README.md
@@ -0,0 +1,44 @@
+# Task chat scrollback regression
+
+Run from the repository root:
+
+```sh
+pnpm exec playwright test --config tests/perf/task-chat/playwright.config.ts
+```
+
+The test starts an isolated Vite server on port 4197. It needs no database,
+credentials, or running agents. It renders the real `TaskChatThreadView`, bubbles,
+markdown, run folds, and scroller with 200 long responses and 4,000 historical
+tool entries. A synthetic live tail updates ten times per second. A second case
+recreates the history objects on each update, as transcript projection can do.
+The harness uses a plain reply textarea; it does not test the full task composer
+or server-to-browser transport.
+
+Chromium samples main-thread task time for three seconds. Both cases must stay
+below 50% main-thread utilization and deliver at least 20 updates. The ceiling
+leaves room for machine variation while detecting the original saturation.
+Performance JSON is attached to each test result. The test also checks reading
+position during updates, typing, return to latest, lazy tool inspection, and
+preserved tool expansion across closing/reopening a run.
+
+For manual inspection, start the same server and open
+`http://127.0.0.1:4197/tests/task-chat-perf.html`:
+
+```sh
+pnpm --filter @paperclipai/ui exec vite --host 127.0.0.1 --port 4197 --strictPort
+```
+
+This HTML entry is not part of the shipped application build.
+
+## Reproduction and results
+
+On macOS Chromium against commit `a05b828bc`, the original code consumed 97.8%
+of the main thread and delivered 8 updates. With the fix, the same dev fixture
+consumed 11.6% for tail-only updates and 31.5% for recreated history, delivering
+32 updates in both cases. These are local samples, not cross-machine promises.
+
+The normal Vitest suite includes deterministic guards in
+`TaskChatThreadView.performance.test.tsx` and `TaskChatTurn.test.tsx`: unchanged
+history does not rerender or reparse markdown, changed content/gallery callbacks
+remain current, and collapsed tools mount on demand. These run in ordinary CI;
+the browser performance test is opt-in.
diff --git a/tests/perf/task-chat/playwright.config.ts b/tests/perf/task-chat/playwright.config.ts
new file mode 100644
index 0000000000..563486b1b8
--- /dev/null
+++ b/tests/perf/task-chat/playwright.config.ts
@@ -0,0 +1,15 @@
+import { defineConfig } from "@playwright/test";
+
+export default defineConfig({
+ testDir: ".",
+ testMatch: "*.spec.ts",
+ workers: 1,
+ timeout: 120_000,
+ use: { baseURL: "http://127.0.0.1:4197", viewport: { width: 1440, height: 900 }, trace: "retain-on-failure" },
+ webServer: {
+ command: "pnpm --filter @paperclipai/ui exec vite --host 127.0.0.1 --port 4197 --strictPort",
+ url: "http://127.0.0.1:4197/tests/task-chat-perf.html",
+ reuseExistingServer: !process.env.CI,
+ timeout: 120_000,
+ },
+});
diff --git a/tests/perf/task-chat/scrollback.spec.ts b/tests/perf/task-chat/scrollback.spec.ts
new file mode 100644
index 0000000000..6e0dec9fea
--- /dev/null
+++ b/tests/perf/task-chat/scrollback.spec.ts
@@ -0,0 +1,56 @@
+import { expect, test } from "@playwright/test";
+
+for (const reproject of [false, true]) {
+test(`long scrollback stays responsive (${reproject ? "reprojected history" : "tail only"})`, async ({ page }, testInfo) => {
+ const errors: string[] = [];
+ page.on("pageerror", (error) => errors.push(error.message));
+ await page.goto("/tests/task-chat-perf.html");
+ await expect(page.getByTestId("task-chat-scroller")).toBeVisible({ timeout: 90_000 });
+ await expect(page.locator('[data-thread-anchor^="history-"]')).toHaveCount(200);
+ if (reproject) await page.getByLabel("Recreate history objects").check();
+ const session = await page.context().newCDPSession(page);
+ await session.send("Performance.enable");
+ const read = async () => Object.fromEntries((await session.send("Performance.getMetrics")).metrics.map(({ name, value }) => [name, value]));
+ await page.getByRole("button", { name: "Start streaming" }).click();
+ const before = await read();
+ await page.waitForTimeout(3000);
+ const after = await read();
+ const metrics = {
+ mainThreadBusyPercent: 100 * (after.TaskDuration - before.TaskDuration) / (after.Timestamp - before.Timestamp),
+ scriptMs: 1000 * (after.ScriptDuration - before.ScriptDuration),
+ layoutMs: 1000 * (after.LayoutDuration - before.LayoutDuration),
+ domNodes: await page.locator("*").count(),
+ ticks: Number(await page.getByTestId("stream-tick").textContent()),
+ };
+ console.log(JSON.stringify(metrics));
+ await testInfo.attach("performance.json", { body: JSON.stringify(metrics, null, 2), contentType: "application/json" });
+ // Read scrollback without getting pulled down by ongoing live updates.
+ const scroller = page.getByTestId("task-chat-scroller");
+ await scroller.hover();
+ await page.mouse.wheel(0, -700);
+ await expect(page.getByRole("button", { name: "Scroll to latest" })).toBeVisible();
+ const top = await scroller.evaluate((element) => element.scrollTop);
+ await page.getByRole("textbox", { name: "Reply" }).fill("Reply remains usable during streaming.");
+ await page.waitForTimeout(300);
+ expect(await scroller.evaluate((element) => element.scrollTop)).toBeCloseTo(top, 0);
+ await expect(page.getByRole("textbox", { name: "Reply" })).toHaveValue("Reply remains usable during streaming.");
+ await page.getByRole("button", { name: "Scroll to latest" }).click();
+ await expect.poll(() => scroller.evaluate((element) => element.scrollHeight - element.clientHeight - element.scrollTop)).toBeLessThan(2);
+ await page.getByRole("button", { name: "Stop streaming" }).click();
+ await expect(page.getByTestId("task-chat-tool-card")).toHaveCount(0);
+ const summary = page.getByTestId("task-chat-turn-summary").last();
+ await summary.click();
+ await expect(summary).toHaveAttribute("aria-expanded", "true");
+ await expect(page.getByTestId("task-chat-tool-card")).toHaveCount(20);
+ const tool = page.getByTestId("task-chat-tool-card").first();
+ await tool.getByRole("button").click();
+ await expect(tool).toContainText("File inspected successfully.");
+ await summary.click();
+ await summary.click();
+ await expect(tool).toContainText("File inspected successfully.");
+ expect(errors).toEqual([]);
+ // A broad regression ceiling, not a machine-specific benchmark target.
+ expect(metrics.mainThreadBusyPercent).toBeLessThan(50);
+ expect(metrics.ticks).toBeGreaterThanOrEqual(20);
+});
+}
diff --git a/ui/src/components/TaskChatThread.tsx b/ui/src/components/TaskChatThread.tsx
index 1fff76d2db..b1e5bb8fda 100644
--- a/ui/src/components/TaskChatThread.tsx
+++ b/ui/src/components/TaskChatThread.tsx
@@ -2454,6 +2454,11 @@ export function TaskChatThread(props: TaskChatThreadProps) {
],
);
+ const renderBrief = useCallback(
+ () => issueBrief ? : null,
+ [issueBrief],
+ );
+
const assignedAgentForNotice = useMemo(() => {
if (!currentAssigneeValue?.startsWith("agent:")) return null;
const assigneeAgentId = currentAssigneeValue.slice("agent:".length);
@@ -2707,11 +2712,7 @@ export function TaskChatThread(props: TaskChatThreadProps) {
attachments={attachments}
header={threadHeaderWithBlockers}
renderInteraction={renderInteraction}
- renderBrief={
- issueBrief
- ? () =>
- : undefined
- }
+ renderBrief={renderBrief}
renderMessageActions={renderMessageActions}
renderQueuedAction={renderQueuedAction}
onTryAgainNoLiveExecutionPath={
diff --git a/ui/src/components/task-chat/TaskChatBubble.tsx b/ui/src/components/task-chat/TaskChatBubble.tsx
index cbcc99d9b7..c6abf055a1 100644
--- a/ui/src/components/task-chat/TaskChatBubble.tsx
+++ b/ui/src/components/task-chat/TaskChatBubble.tsx
@@ -1,4 +1,4 @@
-import { useContext, useState, type ReactNode } from "react";
+import { useCallback, useContext, useState, type ReactNode } from "react";
import type { IssueAttachment } from "@paperclipai/shared";
import { IssueGalleryContext } from "@/context/IssueGalleryContext";
import { cn } from "@/lib/utils";
@@ -156,9 +156,11 @@ export function TaskChatBubble({
// Task attachments share the page gallery; standalone images retain the bubble viewer.
const openIssueGallery = useContext(IssueGalleryContext);
const [lightboxSrc, setLightboxSrc] = useState(null);
- const openImage = (src: string) => {
+ // Keep MarkdownBody's memo boundary intact when only the live tail changes.
+ // A fresh callback here reparses every historical response on every update.
+ const openImage = useCallback((src: string) => {
if (!openIssueGallery?.(src)) setLightboxSrc(src);
- };
+ }, [openIssueGallery]);
if (item.interstitial) {
// Interstitial updates are ephemeral (PAP-361): while streaming the text
// lives on the live parent row's line (TaskChatStatusItem.selfTalk), and
diff --git a/ui/src/components/task-chat/TaskChatThreadView.performance.test.tsx b/ui/src/components/task-chat/TaskChatThreadView.performance.test.tsx
new file mode 100644
index 0000000000..e9fed6e648
--- /dev/null
+++ b/ui/src/components/task-chat/TaskChatThreadView.performance.test.tsx
@@ -0,0 +1,77 @@
+// @vitest-environment jsdom
+
+import { memo } from "react";
+import { flushSync } from "react-dom";
+import { createRoot, type Root } from "react-dom/client";
+import { afterEach, beforeEach, expect, it, vi } from "vitest";
+import { TaskChatThreadView } from "./TaskChatThreadView";
+import { IssueGalleryContext } from "@/context/IssueGalleryContext";
+import type { TaskChatItem, TaskChatMessageItem } from "./task-chat-model";
+
+const renders = vi.hoisted(() => vi.fn());
+vi.mock("@/components/MarkdownBody", () => ({
+ // Preserve the real markdown component's memo contract while counting work.
+ MarkdownBody: memo((props: { children: string; onImageClick?: (src: string) => void }) => {
+ renders(props.children);
+ return ;
+ }),
+}));
+
+let host: HTMLDivElement;
+let root: Root;
+beforeEach(() => {
+ renders.mockClear();
+ host = document.body.appendChild(document.createElement("div"));
+ root = createRoot(host);
+});
+afterEach(() => {
+ flushSync(() => root.unmount());
+ host.remove();
+});
+
+const history: TaskChatMessageItem[] = Array.from({ length: 200 }, (_, index) => ({
+ id: `message-${index}`, kind: "message", author: "agent", text: `Response ${index}`,
+}));
+
+it("does no historical render work for a live-tail-only update", () => {
+ const actions = vi.fn(() => null);
+ const render = (tick: number) => flushSync(() => root.render(
+ Live {tick}