From ee532c2647300e136f1a0336eb97701a97461b81 Mon Sep 17 00:00:00 2001 From: Luis Esteban Acevedo Ladino Date: Fri, 24 Apr 2026 09:34:38 -0500 Subject: [PATCH] docs --- docs/agent-tool-specs.md | 451 ++++++++++++++++++ docs/agent-tools.md | 324 +++++++++++++ .../exploration.md | 292 ++++++++++++ .../gemini-analyze-video-fase-1/proposal.md | 62 +++ 4 files changed, 1129 insertions(+) create mode 100644 docs/agent-tool-specs.md create mode 100644 docs/agent-tools.md create mode 100644 openspec/changes/gemini-analyze-video-fase-1/exploration.md create mode 100644 openspec/changes/gemini-analyze-video-fase-1/proposal.md diff --git a/docs/agent-tool-specs.md b/docs/agent-tool-specs.md new file mode 100644 index 00000000..57db4d62 --- /dev/null +++ b/docs/agent-tool-specs.md @@ -0,0 +1,451 @@ +# Agent Tool Specs — NeuralCut + +Este documento define las primitives iniciales del agente. Cada tool debe ser pequeña, accionable y componible. El agente puede resolver features complejas combinando estas tools, pero ninguna tool debe intentar “hacer todo”. + +## Principios generales + +- Las tools reciben datos estructurados, no instrucciones libres. +- Las tools deben validar inputs y devolver errores claros. +- Las tools que modifican el editor deben usar los comandos/managers existentes para preservar undo/redo cuando aplique. +- El agente debe usar `list_project_assets` y `list_timeline` antes de editar cuando no tenga IDs concretos. +- Gemini se usa para comprensión multimodal; la edición real se ejecuta con tools determinísticas. + +--- + +## 1. `list_project_assets` + +### Propósito +Listar assets disponibles en el proyecto para que el agente sepa qué archivos existen y cuáles están usados en el timeline. + +### Input +```ts +{ + filter?: "all" | "used" | "unused"; + type?: "all" | "video" | "audio" | "image"; +} +``` + +### Output +```ts +{ + assets: Array<{ + id: string; + name: string; + type: "video" | "audio" | "image"; + duration?: number; + usedInTimeline: boolean; + }>; +} +``` + +### Requirements +- MUST return all loaded project assets by default. +- MUST support filtering by usage: `used`, `unused`, `all`. +- MUST support filtering by asset type. +- MUST include stable internal `id` values. +- MUST NOT mutate editor state. + +### Errors +- If no project is loaded, return `{ error: "No active project" }`. + +--- + +## 2. `list_timeline` + +### Propósito +Listar el estado actual del timeline para que el agente pueda referenciar clips, textos, stickers y otros elementos por `trackId`/`elementId`. + +### Input +```ts +{} +``` + +### Output +```ts +{ + tracks: Array<{ + trackId: string; + type: "main" | "overlay" | "audio" | "text" | "effect"; + elements: Array<{ + elementId: string; + type: string; + assetId?: string; + name?: string; + start: number; + end: number; + }>; + }>; +} +``` + +### Requirements +- MUST return a structured timeline summary. +- MUST include `trackId` and `elementId` for editable elements. +- MUST include timing in seconds or a clearly documented unit. +- MUST NOT mutate editor state. + +### Errors +- If no active scene/timeline exists, return `{ error: "No active timeline" }`. + +--- + +## 3. `load_asset_context` + +### Propósito +Cargar un asset en el contexto multimodal del agente usando Gemini. Esta es la base para que el agente entienda video/audio/imagen antes de editar. + +### Input +```ts +{ + assetId: string; +} +``` + +### Output +```ts +{ + assetId: string; + status: "loaded" | "processing"; + cached: boolean; + provider: "gemini"; + fileUri?: string; + summary?: string; +} +``` + +### Requirements +- MUST resolve the asset by internal `assetId`. +- MAY support fallback by filename if unambiguous, but internal ID is preferred. +- MUST upload/process the asset with Gemini when not cached. +- MUST cache the provider file reference by `assetId` to avoid repeated uploads. +- MUST keep API keys server-side. +- MUST support video first; audio/image support may be added if Gemini path supports it cleanly. +- MUST NOT edit the timeline. + +### Cache behavior +- Cache is application/orchestrator infrastructure, not necessarily a visible LLM tool. +- Cache should store at least: + ```ts + { + assetId: string; + provider: "gemini"; + fileUri: string; + status: "active" | "processing" | "failed"; + createdAt: number; + } + ``` + +### Errors +- Unknown asset: `{ error: "Asset not found" }`. +- Unsupported type: `{ error: "Unsupported asset type" }`. +- Upload/processing failure: `{ error: "Failed to load asset context" }` with safe details. + +### Non-goals +- No editing actions. +- No `@asset` UI autocomplete. +- No streaming progress in the first implementation. + +--- + +## 4. `cut_segment` + +### Propósito +Cortar o remover un rango de tiempo del timeline. + +### Input +```ts +{ + start: number; + end: number; + mode: "remove" | "keep"; +} +``` + +### Output +```ts +{ + success: boolean; + affectedElements: string[]; +} +``` + +### Requirements +- MUST validate `start < end`. +- MUST operate on timeline time, using the editor’s canonical time unit internally. +- MUST use existing timeline/command infrastructure when possible. +- `remove` MUST remove the selected time range. +- `keep` MUST preserve only the selected time range and remove outside material, if feasible in current editor model. +- MUST preserve undo/redo behavior if the editor supports it for the operation. + +### Errors +- Invalid range: `{ error: "Invalid time range" }`. +- Empty timeline: `{ error: "No timeline content" }`. +- Unsupported mode: `{ error: "Unsupported cut mode" }`. + +--- + +## 5. `add_text` + +### Propósito +Agregar texto visual al timeline. Esta primitive cubre títulos, hooks, labels y subtítulos básicos. + +### Input +```ts +{ + text: string; + start: number; + end: number; + position: "top" | "center" | "bottom"; + style?: "plain" | "subtitle" | "hook" | "label"; +} +``` + +### Output +```ts +{ + elementId: string; + trackId: string; +} +``` + +### Requirements +- MUST validate non-empty `text`. +- MUST validate `start < end`. +- MUST create a text element using existing timeline APIs. +- MUST map `position` to a sensible default placement. +- SHOULD provide simple style presets, but implementation may start with defaults. +- MUST NOT generate full subtitles automatically; that is a flow using transcript/context + repeated `add_text` calls. + +### Errors +- Empty text: `{ error: "Text is required" }`. +- Invalid range: `{ error: "Invalid time range" }`. + +--- + +## 6. `update_text` + +### Propósito +Editar un texto existente en el timeline. + +### Input +```ts +{ + trackId: string; + elementId: string; + text?: string; + start?: number; + end?: number; + position?: "top" | "center" | "bottom"; +} +``` + +### Output +```ts +{ + success: boolean; + elementId: string; +} +``` + +### Requirements +- MUST find an existing text element by `trackId` + `elementId`. +- MUST only update provided fields. +- MUST validate timing if `start`/`end` are provided. +- MUST preserve undo/redo behavior if supported. + +### Errors +- Missing element: `{ error: "Text element not found" }`. +- Wrong element type: `{ error: "Element is not text" }`. +- Invalid range: `{ error: "Invalid time range" }`. + +--- + +## 7. `add_media_to_timeline` + +### Propósito +Agregar un asset existente al timeline. + +### Input +```ts +{ + assetId: string; + startTime: number; + trackType: "main" | "overlay" | "audio"; +} +``` + +### Output +```ts +{ + elementId: string; + trackId: string; +} +``` + +### Requirements +- MUST resolve asset by `assetId`. +- MUST validate asset type compatibility with `trackType`. +- MUST insert the element using existing timeline APIs. +- SHOULD choose a sensible track when one is not obvious, but first version requires explicit `trackType`. + +### Errors +- Asset not found: `{ error: "Asset not found" }`. +- Invalid track type: `{ error: "Invalid track type for asset" }`. + +--- + +## 8. `delete_element` + +### Propósito +Eliminar un elemento específico del timeline. + +### Input +```ts +{ + trackId: string; + elementId: string; +} +``` + +### Output +```ts +{ + success: boolean; +} +``` + +### Requirements +- MUST find element by `trackId` + `elementId`. +- MUST remove only that element. +- MUST preserve undo/redo behavior if supported. + +### Errors +- Missing element: `{ error: "Element not found" }`. + +--- + +## 9. `set_volume` + +### Propósito +Ajustar volumen de un elemento de audio o video. + +### Input +```ts +{ + trackId: string; + elementId: string; + volume: number; +} +``` + +### Output +```ts +{ + success: boolean; + volume: number; +} +``` + +### Requirements +- MUST validate `volume` in range `0..1`. +- MUST only apply to elements that support audio volume. +- MUST preserve undo/redo behavior if supported. + +### Errors +- Invalid volume: `{ error: "Volume must be between 0 and 1" }`. +- Unsupported element: `{ error: "Element does not support volume" }`. + +--- + +## 10. `add_sticker` + +### Propósito +Agregar un sticker existente al timeline. + +### Input +```ts +{ + stickerId: string; + start: number; + end: number; + position: "top-left" | "top-right" | "center" | "bottom-left" | "bottom-right"; +} +``` + +### Output +```ts +{ + elementId: string; + trackId: string; +} +``` + +### Requirements +- MUST resolve sticker by `stickerId`. +- MUST validate `start < end`. +- MUST create a timeline element using existing sticker/timeline APIs. +- MUST map `position` to sensible coordinates. + +### Errors +- Sticker not found: `{ error: "Sticker not found" }`. +- Invalid range: `{ error: "Invalid time range" }`. + +--- + +## 11. `apply_effect` + +### Propósito +Aplicar un efecto existente a un clip. En el estado actual del repo, el efecto real disponible parece ser `blur`. + +### Input +```ts +{ + trackId: string; + elementId: string; + effectType: "blur"; + params?: { + intensity?: number; + }; +} +``` + +### Output +```ts +{ + effectId: string; + elementId: string; +} +``` + +### Requirements +- MUST validate the element is visual and supports effects. +- MUST validate `effectType` exists in the effects registry. +- MUST apply default params when `params` are omitted. +- MUST validate `intensity` if provided. +- MUST preserve undo/redo behavior if supported. + +### Errors +- Effect not found: `{ error: "Effect not found" }`. +- Unsupported element: `{ error: "Element does not support effects" }`. +- Invalid params: `{ error: "Invalid effect parameters" }`. + +--- + +## Existing/secondary tool: `transcribe_video` + +### Propósito +Transcribir audio usando Whisper local. Es útil para captions y edición por texto, pero no debe ser la primitive principal de comprensión si `load_asset_context` con Gemini está disponible. + +### Input +```ts +{ + assetId?: string; + language?: string; + modelId?: string; +} +``` + +### Estado +- Implementada. +- Debe mantenerse como secundaria. +- Puede ser usada por flujos de subtítulos. diff --git a/docs/agent-tools.md b/docs/agent-tools.md new file mode 100644 index 00000000..85a3f506 --- /dev/null +++ b/docs/agent-tools.md @@ -0,0 +1,324 @@ +# Agent Tools — Lista propuesta + +Este documento resume las tools propuestas para el agente de NeuralCut. La idea es mantenerlas **primitivas, acotadas y componibles**, no crear una tool por cada feature de marketing. + +## Principio + +El agente debe combinar primitives simples: + +- primero entiende el proyecto/assets, +- luego decide qué acción hacer, +- después ejecuta operaciones concretas del editor. + +No queremos tools gigantes tipo “hazme un reel completo”. Eso debe ser un flujo del agente usando varias tools pequeñas. + +--- + +## Tools de contexto y percepción + +### `list_project_assets` + +Lista assets conocidos del proyecto. + +```ts +{ + filter?: "all" | "used" | "unused"; + type?: "all" | "video" | "audio" | "image"; +} +``` + +Uso: +- saber qué archivos existen, +- distinguir assets usados/no usados, +- decidir qué asset cargar o editar. + +--- + +### `list_timeline` + +Lista el estado actual del timeline. + +```ts +{} +``` + +Uso: +- saber qué clips/textos/stickers/effects existen, +- obtener `trackId` y `elementId`, +- preparar operaciones como cortar, borrar o editar. + +--- + +### `load_asset_context` + +Carga un asset en el contexto multimodal del agente usando Gemini. + +```ts +{ + assetId: string; +} +``` + +Uso: +- subir/procesar video, audio o imagen con Gemini, +- dejar el asset disponible para razonamiento posterior, +- cachear la referencia para no subirlo repetidamente. + +> Esta reemplaza la idea anterior de `analyze_video({ question })` o `analyze_asset({ analysisType })`. La tool no “pregunta”; solo carga el asset al contexto. + +--- + +### `get_asset_context` *(infra interna, no tool visible)* + +Recupera información/cache de un asset ya cargado. + +```ts +{ + assetId: string; +} +``` + +Uso: +- reutilizar análisis/contexto previo, +- evitar re-upload, +- permitir que el agente sepa si un asset ya está disponible para Gemini. + +> Esta NO debería exponerse necesariamente como tool al LLM. Es más sano tratarla como infraestructura interna del agente/orquestador: `load_asset_context` carga y cachea; el orquestador/store puede consultar el cache e inyectar ese estado en el prompt/contexto sin hacer que el modelo llame una tool administrativa. + +--- + +## Tools de edición primitivas + +### `cut_segment` + +Corta o remueve un rango de tiempo del timeline. + +```ts +{ + start: number; + end: number; + mode: "remove" | "keep"; +} +``` + +Uso: +- quitar silencios, +- quitar errores, +- recortar intro/outro, +- construir highlights. + +--- + +### `add_media_to_timeline` + +Agrega un asset existente al timeline. + +```ts +{ + assetId: string; + startTime: number; + trackType: "main" | "overlay" | "audio"; +} +``` + +Uso: +- insertar clips, +- agregar música, +- poner b-roll, +- agregar imágenes/logos. + +--- + +### `delete_element` + +Elimina un elemento específico del timeline. + +```ts +{ + trackId: string; + elementId: string; +} +``` + +Uso: +- borrar clips, +- borrar textos, +- borrar stickers, +- borrar efectos standalone si aplica. + +--- + +### `add_text` + +Agrega texto visual al timeline. + +```ts +{ + text: string; + start: number; + end: number; + position: "top" | "center" | "bottom"; + style?: "plain" | "subtitle" | "hook" | "label"; +} +``` + +Uso: +- subtítulos, +- hooks, +- títulos, +- labels, +- texto explicativo. + +--- + +### `update_text` + +Actualiza un texto existente. + +```ts +{ + trackId: string; + elementId: string; + text?: string; + start?: number; + end?: number; + position?: "top" | "center" | "bottom"; +} +``` + +Uso: +- corregir texto, +- cambiar timing, +- mover captions/hooks. + +--- + +### `add_sticker` + +Agrega un sticker al timeline. + +```ts +{ + stickerId: string; + start: number; + end: number; + position: "top-left" | "top-right" | "center" | "bottom-left" | "bottom-right"; +} +``` + +Uso: +- agregar elementos visuales simples, +- reacciones, +- énfasis gráfico. + +--- + +### `set_volume` + +Ajusta volumen de un elemento de audio/video. + +```ts +{ + trackId: string; + elementId: string; + volume: number; // 0–1 +} +``` + +Uso: +- bajar música, +- subir voz, +- mutear clip, +- balance básico. + +--- + +### `apply_effect` + +Aplica un efecto existente a un clip. + +```ts +{ + trackId: string; + elementId: string; + effectType: "blur"; + params?: { + intensity?: number; + }; +} +``` + +Uso: +- aplicar blur. + +> El repo actualmente tiene infraestructura de efectos, pero el efecto real registrado parece ser solo `blur`. Corrección de color/LUTs todavía no está disponible. + +--- + +## Tools existentes o en standby + +### `transcribe_video` + +Transcribe audio de video/audio usando Whisper local. + +```ts +{ + assetId?: string; + language?: string; + modelId?: string; +} +``` + +Uso: +- subtítulos, +- captions, +- búsqueda de frases exactas, +- edición por texto. + +Estado: +- implementada, +- útil como herramienta secundaria, +- no debe ser la tool principal para entender el video si Gemini puede cargar el asset multimodalmente. + +--- + +## Ideas descartadas o diferidas + +### `analyze_video({ question })` + +Descartada como primitive principal porque mete “chat dentro de la tool”. Mejor usar `load_asset_context` y dejar que el agente razone con el contexto cargado. + +### `analyze_asset({ analysisType })` + +Descartada porque `analysisType` introduce categorías artificiales. Gemini debe cargar el asset completo al contexto; el agente decide qué hacer después. + +### `remove_silences` + +Diferida/no prioritaria porque se puede expresar como flujo con `cut_segment`. + +### `generate_subtitles` + +Diferida porque se puede construir con `transcribe_video` + `add_text`. + +### `color_correct` / `apply_lut` + +Diferida porque el repo actual no parece tener corrección de color/LUTs implementados todavía. + +--- + +## Orden recomendado + +1. `list_project_assets` +2. `list_timeline` +3. `load_asset_context` +4. `cut_segment` +5. `add_text` +6. `add_media_to_timeline` +7. `delete_element` +8. `set_volume` +9. `add_sticker` +10. `apply_effect` + +Infra interna asociada: + +- `get_asset_context` / cache lookup para `assetId -> Gemini fileUri/status/summary`, no necesariamente visible para el LLM. diff --git a/openspec/changes/gemini-analyze-video-fase-1/exploration.md b/openspec/changes/gemini-analyze-video-fase-1/exploration.md new file mode 100644 index 00000000..8429f8fe --- /dev/null +++ b/openspec/changes/gemini-analyze-video-fase-1/exploration.md @@ -0,0 +1,292 @@ +# Exploration: `analyze_video` — Gemini Video Understanding (Phase 1) + +## Current State + +### Shipped infrastructure + +Three completed changes underpin this work: + +| Change | What shipped | +|--------|-------------| +| `infra-habilitadora-fase-1` | Full agent pipeline: orchestrator (client-side), tool registry, mock tool, context adapter, system prompt, chat UI | +| `gemini-adapter-fase-1` | `GeminiAdapter` behind `ProviderAdapter` seam — text-only `chat()` via `@google/generative-ai` v0.24.0 | +| `transcribe-video-fase-1` | `transcribe_video` tool — Whisper-based, client-side execution, returns `TranscriptionToolResult` with segments + timestamps. Fully implemented and verified. | + +### How the agent pipeline works today + +``` +User message → ChatPanel → orchestrator.run(messages, context) + → POST /api/agent/chat (server) + → createProvider(config) → GeminiAdapter.chat({ messages, systemPrompt, tools }) + → model.generateContent({ contents, systemInstruction, tools }) + → { content, toolCalls? } + → resolveToolCalls(toolCalls, context) (client) + → toolRegistry.get(name).execute(args, context) + → ToolResult → appended as tool_result message + → Loop back to POST (max 8 iterations) +``` + +**Key architectural facts**: +- **Server route** (`route.ts`): Stateless, holds `LLM_API_KEY`, creates adapter, calls `chat()`. Exposes tool schemas to the LLM but NEVER executes tools. +- **Client orchestrator**: Holds `AgentContext`, resolves tool calls, executes tools in-browser where `EditorCore` is accessible. +- **Tools are client-side only**: `ToolDefinition.execute()` runs in the browser, has access to `EditorContextAdapter.resolveAssetFile()` and `EditorCore`. +- **Tool schemas are duplicated**: `route.ts` has `providerToolDefs` (stub `execute`) for the LLM; `toolRegistry` has real definitions for execution. + +### Gemini SDK already installed + +`@google/generative-ai` v0.24.0 is in `apps/web/package.json`. The `GeminiAdapter` uses `GoogleGenerativeAI` → `getGenerativeModel()` → `generateContent()`. The SDK also exports: + +- **`GoogleAIFileManager`** from `@google/generative-ai/server` — server-side file upload/management +- **`FileState`** enum — `PROCESSING`, `ACTIVE`, `FAILED` +- **`FileDataPart`** — `{ fileData: { fileUri, mimeType } }` content part for `generateContent()` + +### EditorContextAdapter + +Already provides the sanctioned path for tools to obtain a `File`: + +```ts +EditorContextAdapter.resolveAssetFile(assetId?: string): File | null +EditorContextAdapter.getAssetHasAudio(assetId: string): boolean | undefined +``` + +--- + +## Affected Areas + +| File | Why affected | +|------|-------------| +| `apps/web/src/agent/tools/analyze-video.tool.ts` | **NEW** — `analyze_video` tool definition (client-side) | +| `apps/web/src/app/api/agent/analyze-video/route.ts` | **NEW** — server route: upload video to Gemini File API + poll + query | +| `apps/web/src/agent/orchestrator.ts` | Add side-effect import of new tool | +| `apps/web/src/app/api/agent/chat/route.ts` | Add `analyze_video` to `providerToolDefs`, optionally remove/demote `transcribe_video` | +| `apps/web/src/agent/system-prompt.ts` | May update guidance to prefer `analyze_video` | +| `apps/web/src/components/editor/panels/chat/message-bubble.tsx` | May add rendering for `AnalyzeVideoResult` shape | + +**Unchanged** (confirmed safe): +- `providers/gemini.ts` — adapter handles text chat; video goes through a separate route using `GoogleAIFileManager` +- `providers/types.ts` — `ProviderAdapter` contract unchanged +- `agent/types.ts` — `AgentContext`, `ToolDefinition` unchanged +- `agent/context.ts` — `resolveAssetFile()` already returns `File` +- `tools/transcribe-video.tool.ts` — stays registered, just de-emphasized in tool list + +--- + +## Critical Constraint: File API Is Server-Side Only + +The Gemini Files API requires the API key and can only be called server-side. The current architecture has tools executing **client-side**. This means `analyze_video` cannot be a pure client-side tool like `transcribe_video`. It needs a **server route** to handle the upload-to-Gemini flow. + +### Proposed flow + +``` +LLM decides: analyze_video({ assetId: "v1", question: "What happens in this video?" }) + → Client orchestrator resolves tool call + → tool.execute(args, context) + → EditorContextAdapter.resolveAssetFile("v1") → File blob + → POST /api/agent/analyze-video (FormData: video file + question + assetId) + → Server: GoogleAIFileManager.uploadFile(buffer, { mimeType, displayName }) + → Server: poll file.state until ACTIVE (or timeout) + → Server: model.generateContent([ question, { fileData: { fileUri, mimeType } }]) + → Server: return { answer, duration, ... } + → Tool returns structured AnalyzeVideoResult +``` + +### Why not inline base64? + +Gemini supports `InlineDataPart` with base64-encoded video, but: +- **Size limit**: Inline data is practical only for files under ~20MB raw (~27MB base64). Real video projects regularly exceed this. +- **No reuse**: The file is sent every call. The File API gives a `fileUri` that persists (up to 48h free / 7d paid), enabling reuse. +- **Processing state**: The File API handles video decoding, audio extraction, and indexing server-side. Inline data forces the model to process on every call. + +**Verdict**: Use the File API for production readiness. Phase 1 uses the File API exclusively. + +--- + +## Approaches + +### 1. Dedicated `/api/agent/analyze-video` Route (Recommended) + +Create a new server route that handles the full upload→poll→query flow. The client-side tool sends the video file as FormData, the server does everything Gemini-side, and returns a structured answer. + +- **Pros**: + - Clean separation: upload+analyze logic is entirely server-side (no API key leakage) + - Reuses existing `GeminiAdapter` pattern (same SDK, same config) + - The tool on the client side is thin — just FormData construction and response parsing + - Natural place for future caching (fileUri by assetId) + - Can report processing progress via streaming in Phase 2 +- **Cons**: + - New route to maintain + - Large FormData upload adds latency (video files can be hundreds of MB) + - Polling loop blocks the server request (video processing can take 30s-5min for long videos) +- **Effort**: Medium + +### 2. Extend `/api/agent/chat` Route for Multimodal + +Add multimodal content support to the existing chat route. The client sends video content parts alongside text messages. The server handles upload transparently within the adapter. + +- **Pros**: + - Single route, single contract + - More "correct" architecturally — the chat route becomes multimodal +- **Cons**: + - **Major refactor of the adapter contract** — `ProviderAdapter.chat()` currently takes text-only `ChatMessage[]`. Adding binary content parts requires interface changes to `ChatMessage`, `ProviderAdapter`, and both adapters. + - Breaks the clean text-only abstraction that `route.ts` depends on + - The upload/poll cycle doesn't fit neatly into the synchronous `chat()` call + - Cross-contaminates text chat and file upload concerns +- **Effort**: High + +### 3. Upload-Then-Chat Two-Step + +Split into two routes: `POST /api/agent/upload-video` (upload + poll → return fileUri) and `POST /api/agent/chat` (modified to accept fileUri references). The client tool orchestrates both calls. + +- **Pros**: + - Separation of upload and query concerns + - fileUri can be cached client-side and reused across queries without re-uploading +- **Cons**: + - Two round trips for a single tool call (upload, then query) + - Client must manage fileUri lifecycle (when to re-upload) + - More complex tool execution logic + - The chat route still needs multimodal content support (same con as Approach 2) +- **Effort**: Medium-High + +--- + +## Recommendation + +**Approach 1: Dedicated `/api/agent/analyze-video` route.** + +Rationale: +1. **Minimal contract changes**: No changes to `ProviderAdapter`, `ChatMessage`, or the existing chat route. The new route is self-contained. +2. **API key safety**: All Gemini interaction happens server-side. The client tool only sends a FormData blob. +3. **Matches the existing pattern**: Just as `transcribe_video` is a client tool that calls services, `analyze_video` is a client tool that calls a server endpoint. The server endpoint is the Gemini File API equivalent of the transcription service. +4. **Cache-friendly**: The route can return the `fileUri` in the response, and in Phase 2, we add a KV cache (Upstash Redis is already installed) keyed by `assetId` to skip re-uploads. +5. **Thin scope**: The tool is `analyze_video({ assetId?, question })` → answer. No editing actions, no timeline mutations. + +--- + +## Recommended Tool Design + +```typescript +// analyze-video.tool.ts + +interface AnalyzeVideoArgs { + assetId?: string; + question: string; +} + +interface AnalyzeVideoResult { + assetName: string; + question: string; + answer: string; + duration: number; +} + +// OR error: +interface AnalyzeVideoError { + error: string; +} +``` + +### Server route contract + +``` +POST /api/agent/analyze-video +Content-Type: multipart/form-data + +Fields: + - video: File (the video blob) + - question: string + - assetId: string (for display/logging) + - mimeType: string (e.g., "video/mp4") + +Response: + { answer: string, duration: number } + OR { error: string } +``` + +### Processing on the server + +``` +1. Parse FormData → extract video buffer + question + mimeType +2. GoogleAIFileManager.uploadFile(buffer, { mimeType, displayName: assetId }) +3. Poll file.state every 5s until ACTIVE (max 120s timeout) +4. model.generateContent([ + { text: question }, + { fileData: { fileUri: file.uri, mimeType: file.mimeType } } + ]) +5. Return { answer: response.text(), duration } +``` + +### Asset resolution + +Reuses the same pattern as `transcribe_video`: +1. If `assetId` provided → resolve that specific asset +2. If only one video/audio asset → use it +3. If multiple → return error asking user to specify + +### De-emphasizing `transcribe_video` + +In `route.ts`, replace `transcribe_video` in `providerToolDefs` with `analyze_video`. The tool stays registered in the client-side registry (available if needed) but the LLM no longer sees it in its tool list. + +```typescript +// route.ts — updated providerToolDefs +const providerToolDefs: ToolDefinition[] = [ + { + name: "analyze_video", + description: "Analyzes a video using Gemini multimodal understanding. Can answer questions about visual content, audio, speech, and overall composition. Returns a structured answer.", + parameters: [ + { key: "assetId", type: "string", required: false }, + { key: "question", type: "string", required: true }, + ], + execute: async () => ({}), // stub — never called server-side + }, +]; +``` + +--- + +## Caching Considerations (Phase 1 vs Phase 2) + +### Phase 1 (this change): No caching +- Every `analyze_video` call uploads the video fresh +- Simple, predictable, no state management +- Downside: redundant uploads for repeated questions about the same video + +### Phase 2: File URI cache via Upstash Redis +- `Upstash Redis` is already a project dependency (`@upstash/redis`, `@upstash/ratelimit`) +- Key: `gemini-file:{sha256hash}` or `gemini-file:{assetId}:{fileSize}` +- Value: `{ fileUri, mimeType, expiresAt }` +- On each call: check cache → if valid, skip upload → use cached fileUri +- Need to handle: file mutation (re-upload if asset changes), TTL alignment with Gemini's expiration + +--- + +## Risks + +1. **Upload size and latency**: Video files can be 100MB+. Uploading to the Gemini File API from a server route adds significant latency. The Next.js route has a default body size limit that may need configuration. **Mitigation**: Phase 1 sets a reasonable max file size (e.g., 500MB). Configure `bodySizeLimit` in the Next.js route config. + +2. **Processing polling blocks the request**: After upload, Gemini needs to process the video (decode frames, extract audio, index). For long videos this can take minutes. The server route holds open during polling. **Mitigation**: Set a polling timeout (e.g., 120s). Return a clear error if processing exceeds the limit. Phase 2 can add async processing with a status check endpoint. + +3. **SDK uploadFile takes file path, not Buffer**: The `GoogleAIFileManager.uploadFile()` API officially takes a `filePath: string`. The docs mention `Buffer` support in the API review, but the TypeScript types may not reflect this. **Mitigation**: If `uploadFile` doesn't accept Buffer directly, use the raw Gemini REST API (`POST /upload/v1beta/files`) with `fetch` for the upload, then use the SDK for `generateContent`. OR write to `/tmp` (works in Node.js but not ideal in serverless). + +4. **`@google/generative-ai` is "deprecated"**: Context7 marks this SDK as deprecated, superseded by `@google/genai`. The newer SDK may have better File API support. **Mitigation**: Phase 1 ships with `@google/generative-ai` (already installed). Note SDK migration as Phase 2 tech debt. The File API behavior is identical across SDKs — only the client library wrapper differs. + +5. **Model availability**: Video understanding requires Gemini 2.5 Flash or Pro (or 1.5 Pro). If the user's `LLM_MODEL` is set to a text-only model, the analyze route needs a separate model config. **Mitigation**: The analyze route reads `LLM_MODEL` but can fall back to `gemini-2.5-flash` if the configured model doesn't support multimodal. Or add a separate `GEMINI_VIDEO_MODEL` env var. + +6. **Serverless timeout**: Next.js on Cloudflare (the deployment target) has request timeout limits. Video processing may exceed these. **Mitigation**: Phase 1 targets local development first. For Cloudflare deployment, use a background worker or queue system in Phase 2. + +7. **API key exposure surface**: The new route handles file uploads and must validate inputs carefully to prevent abuse. **Mitigation**: Use the existing `@upstash/ratelimit` for rate limiting. Validate file MIME type server-side. + +8. **No structured timestamps in Phase 1**: Gemini's response about a video includes natural language time references ("at 0:45 the camera pans") but not structured timestamp objects. **Mitigation**: Accept natural language in Phase 1. Phase 2 can add prompt engineering or response parsing for structured timestamps. + +--- + +## Ready for Proposal + +**Yes.** The exploration identifies: +- One new client-side tool file (`analyze-video.tool.ts`) +- One new server route (`/api/agent/analyze-video/route.ts`) +- Minor wiring changes (orchestrator import, route tool defs update) +- Clear de-emphasis path for `transcribe_video` +- No changes to `ProviderAdapter`, `ChatMessage`, or existing adapter code +- Clear Phase 2 boundary (caching, structured timestamps, progress reporting, async processing) + +**Next**: Run `sdd-propose` with this exploration as input. diff --git a/openspec/changes/gemini-analyze-video-fase-1/proposal.md b/openspec/changes/gemini-analyze-video-fase-1/proposal.md new file mode 100644 index 00000000..b0911042 --- /dev/null +++ b/openspec/changes/gemini-analyze-video-fase-1/proposal.md @@ -0,0 +1,62 @@ +# Proposal: Gemini Analyze Video Phase 1 + +## Intent + +Make Gemini video understanding the primary path for video questions so the agent can inspect the actual selected video, not just a transcript. Keep the API key server-side and avoid widening the existing provider/chat contract. + +## Scope + +### In Scope +- Add tool `analyze_video` with args `{ assetId?: string, question: string }`. +- Send the resolved video file to a new server route that uploads/polls Gemini Files API and returns a natural-language answer with optional timestamp hints. +- Expose `analyze_video` as the primary video-understanding tool and de-emphasize `transcribe_video` in LLM-visible tool definitions/prompting. + +### Out of Scope +- Editing actions or timeline mutations. +- `@asset` UI, streaming/progress UX, and robust caching for large videos. + +## Capabilities + +### New Capabilities +- None + +### Modified Capabilities +- `agent-session-shell`: tool registry/API proxy behavior now prioritizes `analyze_video` for video understanding instead of `transcribe_video` as the primary exposed tool. +- `agent-context-bridge`: system prompt guidance must describe `analyze_video` as the preferred tool for questions about loaded video assets. + +## Approach + +Use a dedicated `POST /api/agent/analyze-video` route. The client tool resolves the asset `File`, sends `FormData`, and the server performs Gemini native upload → processing poll → multimodal query. Keep `ProviderAdapter` text-chat only; keep `transcribe_video` registered as standby. + +## Affected Areas + +| Area | Impact | Description | +|------|--------|-------------| +| `apps/web/src/agent/tools/analyze-video.tool.ts` | New | Client tool contract and file resolution | +| `apps/web/src/app/api/agent/analyze-video/route.ts` | New | Server-side Gemini Files API flow | +| `apps/web/src/app/api/agent/chat/route.ts` | Modified | LLM-visible tool definitions prioritize `analyze_video` | +| `apps/web/src/agent/system-prompt.ts` | Modified | Prompt guidance prefers video analysis over transcription | +| `apps/web/src/agent/orchestrator.ts` | Modified | Register new tool | + +## Risks + +| Risk | Likelihood | Mitigation | +|------|------------|------------| +| Large upload / long processing | Med | Timeout, clear errors, phase-2 progress/caching | +| SDK File API friction | Med | Fallback to raw REST upload if needed | +| Wrong model lacks video support | Low | Route-level multimodal-capable model fallback | + +## Rollback Plan + +Remove `analyze_video` wiring, restore `transcribe_video` as the primary exposed tool, and delete the dedicated analyze route. + +## Dependencies + +- Existing Gemini server SDK and server-side `LLM_API_KEY` +- Existing `EditorContextAdapter.resolveAssetFile()` asset resolution path + +## Success Criteria + +- [ ] The LLM can call `analyze_video({ assetId?, question })` and receive a natural-language answer about the real video file. +- [ ] API key remains server-only; no client-side Gemini upload logic is introduced. +- [ ] `transcribe_video` remains available as secondary/standby, not the primary exposed video tool.