Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
72 changes: 35 additions & 37 deletions docs/Features/Face-Analysis.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,53 +37,51 @@ Select an analyzed video clip to see its boxes, landmarks, and anonymous person
labels over the Preview. Existing yellow face markers remain available on the
timeline. The timeline assigns every anonymous person a stable colour: their
sample dots and source-time range bands share that colour. The Analysis tab
combines source-time facts in one workspace. Its compact map aligns scenes,
combines source-time facts in one workspace. Its collapsible compact map aligns scenes,
cuts, speech, people, motion, focus, quality, audio, and text with the
playhead. Below it, one virtualized scene list shows larger speaker face crops,
word-synchronized dialogue, and scene time in compact scene blobs. Clicking a
blob expands the scene-scoped identity tools, facts, descriptions, quality
notices, and transcript in place.

Expanded scenes keep contextual People and Needs Review controls. Clicking a
person filters the virtualized scene list; **Next appearance** advances through
that identity's source-time appearances. Person identities and individual
appearance crops remain draggable correction sources. Dropping one onto another
person merges a false split or moves that appearance. Small yellow detections
that are unsafe to identify are consolidated into short visual tracks without
claiming an identity; their scene-scoped review crops can be clicked to seek or
dragged onto a confirmed person to assign the track manually. These previews
are resolved only for the bounded virtualized window and open scene, stay in a
bounded browser-memory cache, and are also persisted as versioned JPEGs under
playhead. Below it, one virtualized scene list fills the remaining panel height
and shows visible face crops with separate diarized-speaker abbreviations,
word-synchronized dialogue, and scene time in compact scene blobs. Face IDs and
speaker labels remain distinct until an explicit mapping exists. Scene blobs
stay compact and never expand; clicking a word or scene time seeks directly to
that source position.

Person identities and individual appearance crops remain draggable correction
sources inside Analysis settings. Dropping one onto another person merges a
false split or moves that appearance. Small yellow detections that are unsafe
to identify are consolidated into short visual tracks without claiming an
identity; their review crops can be clicked to seek or dragged onto a confirmed
person to assign the track manually. These previews stay in a bounded
browser-memory cache and are also persisted as versioned JPEGs under
`Cache/face-thumbnails` so they load without video seeks after a refresh or
project reopen.

Below the virtualized scene list, the complete detected-person strip restores
the correction surface for source-wide work: representative found-frame crops,
all appearances, person-to-person merge, appearance reassign, and review-track
assignment by drag and drop. It uses the same durable face correction services
as scene cards; no second people state is stored. Face crops are requested only
once their tile is near the viewport. Equal crop requests share the existing
pending/blob cache, so scrolling does not enqueue a fresh video seek for every
rendered card.

The top of the Analysis tab exposes separate action rows for **Focus & Motion**,
**Faces**, **Scene Cuts**, **Transcript**, and **AI Scenes**. Every row can be analyzed,
reanalyzed, retried, or cancelled without clearing the other results. A
Analysis settings contain the complete detected-person strip for source-wide
correction work: representative found-frame crops, all appearances,
person-to-person merge, appearance reassign, and review-track assignment by
drag and drop. It uses the same durable face correction services as scene
cards; no second people state is stored. Face crops are requested only once
their tile is near the viewport. Equal crop requests share the existing
pending/blob cache.

The top of the Analysis tab exposes compact status pills for **Focus & Motion**,
**Faces**, **Cuts**, **Transcript**, and **AI Scenes**. Every pill can analyze,
reanalyze, retry, continue, or cancel without clearing the other results. A
metrics-only pass preserves compatible face observations. A face-only pass
reuses existing focus and motion samples; when no metrics exist yet, it creates
the inexpensive metrics baseline during the same source decode. At normal
Properties widths the actions form one compact three-column grid (with
responsive two- and one-column fallbacks). **Analyze All** creates a coalesced
Properties widths the pills wrap naturally beside **Analyze all** and a compact
settings disclosure. **Analyze All** creates a coalesced
job graph: repeated clicks observe the same run, compatible completed channels
are reused, Focus/Faces and the cut scan serialize on the shared source decoder,
and independent transcript work may proceed in parallel. AI scene descriptions
remain a separate opt-in action because they can incur provider cost and share
visual content externally.

Above those rows, compact **Scope** (`Source`, `Used Ranges`, `Selection`,
`In-Out`) and **Profile** (`Quick`, `Balanced`, `Deep`, `Custom`) controls show
uncached duration, cache reuse, known frame/sample work, relative cost, and a
benchmark-derived time range when one exists. **Quick** (1 fps) and
The settings disclosure separates shared visual-analysis controls from
transcript controls. Compact **Scope** (`Source`, `Used Ranges`, `Selection`,
`In-Out`) and **Profile** (`Quick`, `Balanced`, `Deep`, `Custom`) choices stay
visible without estimate or explanatory rows. **Quick** (1 fps) and
**Balanced** (2 fps) are execution settings for local Focus/Motion and Faces:
the runner receives the selected clipped source range and explicit sample
cadence. Used Ranges, Selection, and overlapping In/Out therefore analyze only
Expand Down Expand Up @@ -216,9 +214,9 @@ to use WebGPU for rendering.

Faces smaller than 36 pixels in the 640-pixel analysis frame are intentionally
excluded from automatic identity grouping. They remain visible as yellow
Preview overlays and as scene-scoped **Needs review** crops in the Analysis
workspace, because title-card grids and background footage are too small for
dependable anonymous matching.
Preview overlays and as **Needs review** crops in Analysis settings, because
title-card grids and background footage are too small for dependable anonymous
matching.

## Models and licenses

Expand Down
61 changes: 33 additions & 28 deletions docs/Features/UI-Panels.md
Original file line number Diff line number Diff line change
Expand Up @@ -374,31 +374,35 @@ same source-handle edges.
| **Masks** | Mask shapes with mode and feather controls |
| **Analysis** | Shared analysis map, transcript controls, and compact scene blobs with faces, synchronized dialogue, cuts, metrics, quality, and descriptions |

The Analysis workspace uses one source-time model for its graph and scene
list. Each virtualized scene blob keeps speaker face crops on the left, larger
The Analysis workspace uses one source-time model for its collapsible graph and
segment list. Visual scene boundaries remain unchanged in the graph, while the
list independently divides long dialogue into readable speech segments. Speaker
changes and long pauses are hard natural boundaries; sentence endings nearest
the 10-second target are preferred, and unpunctuated speech is split at a word
boundary before 15 seconds. This keeps talking-head footage navigable even when
cut detection returns one long scene. The list consumes all remaining panel
height and owns its scroll. Each virtualized segment keeps visible face crops on
the left with a separate transcript-speaker abbreviation beside them, larger
word-synchronized dialogue in the center, and range/duration on the right.
The flat inspector directly beneath the graph replaces the former Current
Frame and Summary boxes with playhead metrics and clip-wide counters. Clicking
a blob expands its scene-scoped people, appearances, identity correction drop
targets, review detections, camera/focus/motion facts, description provenance,
coverage, quality notices, OCR, and transcript in place. Person chips filter
the scene list, while **Next appearance** advances in source time. Face crops
remain lazy because only the bounded virtualized window and the open scene
resolve thumbnails. Contextual People and Needs Review controls stay in their
expanded scene; the source-wide correction strip follows the list.

The scene list is followed by a compact source-wide people/review strip for
the complete correction workflow. It deliberately reuses the same person and
review identities as scene cards: drag a person to merge, an appearance to
move it, or a review track to assign it. Crop loading is viewport-lazy while
the virtualized scene window and crop cache keep scrolling bounded.

The Action Center above the map uses the same flat 2px control treatment as
the rest of Properties. Its analysis actions use three compact equal-height
cards per row at normal panel widths, with integrated status and action
controls instead of full-width button bars. Scope and profile selectors update
a read-only estimate before any work starts, including cached reuse and known
frame/sample counts. A matching real-device benchmark adds a time range.
Speaker labels are not silently equated with anonymous face identities. Segments
remain compact and never expand. Clicking a word seeks to that exact word;
clicking the time seeks to the speech-segment start (or the scene start when no
speech exists). Redundant playhead and clip-summary statistics are omitted. Face
crops remain lazy because only the bounded virtualized window resolves
thumbnails.

The source-wide people/review correction strip lives in Analysis settings so
the normal scene list stays focused. It deliberately reuses the same person
and review identities as scene cards: drag a person to merge, an appearance to
move it, or a review track to assign it. Crop loading remains viewport-lazy.

The Action Center above the map is a single compact row of status pills using
the same flat control treatment as Export. Each pill exposes its channel state
and triggers analyze, reanalyze, continue, retry, or cancel as appropriate.
**Analyze all** and a settings button complete the row. The settings disclosure
uses separate compact areas for shared visual-analysis controls and transcript
controls. It keeps only control titles, scope/profile choices, language/mode,
and actions visible; estimates and explanatory copy do not push the map down.
Quick/Balanced Focus/Motion and Faces execute with the selected source interval
and sampling cadence; frame-accurate cuts remain source-wide. Deep/Custom stay
blocked without qualifying evidence. **Analyze All** deliberately excludes AI
Expand All @@ -407,10 +411,11 @@ provider cost and share visual content externally.

Analysis contains the transcript workspace header for provider and language
selection, start/continue/cancel/clear, fusion progress, coverage, and shared
search. The same search filters the virtualized scene list. Words are grouped
into timestamped speaker turns; clicking a compact or expanded word seeks the
playhead to that source position. During playback, the currently spoken word
and speaker turn are highlighted and the active scene transcript follows them.
search. The same search filters individual rows in the virtualized segment list.
Words are grouped into timestamped speaker turns and readable timed segments;
clicking a word seeks the playhead to that source position. During playback, the
currently spoken word and active speech segment are highlighted and the list
follows them when that segment is visible under the current search filter.
Scrubbing updates the same highlight without animated lag.

Best Quality uses a fixed provider split: Deepgram supplies every displayed word
Expand Down
70 changes: 70 additions & 0 deletions functions/api/kernel/[[path]].ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
import { json, methodNotAllowed } from '../../lib/db';
import type { AppContext, AppRouteHandler } from '../../lib/env';

const DEFAULT_KERNEL_ORIGIN = 'https://fassandra.de';
// Story compiles include provider planning rounds plus the director run.
const FORWARD_TIMEOUT_MS = 240_000;

interface AllowedRoute {
methods: string[];
pattern: RegExp;
requiresUser: boolean;
}

const ALLOWED_ROUTES: AllowedRoute[] = [
{ methods: ['GET'], pattern: /^health$/, requiresUser: false },
{ methods: ['POST'], pattern: /^compile$/, requiresUser: true },
{ methods: ['POST'], pattern: /^runs\/[A-Za-z0-9._:-]+\/complete$/, requiresUser: true },
];

function resolvePath(context: AppContext): string {
const raw = (context.params as Record<string, unknown>).path;
if (Array.isArray(raw)) {
return raw.map(String).join('/');
}
return typeof raw === 'string' ? raw : '';
}

export const onRequest: AppRouteHandler = async (context: AppContext): Promise<Response> => {
const path = resolvePath(context);
const route = ALLOWED_ROUTES.find((candidate) => candidate.pattern.test(path));
if (!route) {
return json({ error: 'Unknown kernel route.' }, { status: 404 });
}
if (!route.methods.includes(context.request.method)) {
return methodNotAllowed(route.methods);
}
if (route.requiresUser && !context.data.user) {
return json({ error: 'Sign in to use the kernel service.' }, { status: 401 });
}

const token = context.env.KERNEL_AUTH_TOKEN?.trim();
if (!token) {
return json({ error: 'Kernel service is not configured.' }, { status: 503 });
}

const origin = context.env.KERNEL_ORIGIN?.trim().replace(/\/+$/, '') || DEFAULT_KERNEL_ORIGIN;
const headers: Record<string, string> = {
Accept: 'application/json',
Authorization: `Bearer ${token}`,
};
const init: RequestInit = {
method: context.request.method,
headers,
signal: AbortSignal.timeout(FORWARD_TIMEOUT_MS),
};
if (context.request.method !== 'GET') {
headers['Content-Type'] = 'application/json';
init.body = await context.request.text();
}

try {
const upstream = await fetch(`${origin}/kernel/${path}`, init);
return new Response(upstream.body, {
status: upstream.status,
headers: { 'Content-Type': 'application/json', 'Cache-Control': 'no-store' },
});
} catch {
return json({ error: 'Kernel service is unreachable.' }, { status: 502 });
}
};
2 changes: 2 additions & 0 deletions functions/lib/env.ts
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,8 @@ export interface Env {
ELEVENLABS_API_KEY?: string;
GOOGLE_CLIENT_ID?: string;
GOOGLE_CLIENT_SECRET?: string;
KERNEL_AUTH_TOKEN?: string;
KERNEL_ORIGIN?: string;
KIEAI_API_KEY?: string;
KIEAI_GENERATION_RATE_LIMITER?: AppDurableObjectNamespace;
KV: AppKVNamespace;
Expand Down
Loading
Loading