From sound to something useful.
Suno turns an uploaded recording into a transcript and a concise summary. The browser handles the upload; a separate worker handles the time-consuming work.
1. Upload without blocking
The Next.js frontend asks FastAPI to create an upload. FastAPI creates a private object key and a multipart upload in AWS S3. The browser uploads up to three 16 MiB parts concurrently using short-lived signed URLs. Each part retries up to three times if the connection fails.
The API checks the part count and total byte size against storage before completing the upload. Large audio never passes through a Next.js request body. The default configuration permits files up to 50 GiB, with no duration cap; available disk space, processing time, and provider credits remain practical limits.
2. A durable background queue
Completing an upload inserts a job and updates the recording in one PostgreSQL transaction. The HTTP request returns immediately. A separate Python worker claims a job using FOR UPDATE SKIP LOCKED, so multiple workers can safely consume the queue.
A worker renews its lease every 15 seconds. If it crashes, another worker can reclaim the job after 180 seconds. Every write checks lease ownership. Transcript sections are checkpointed, so a retry skips sections already saved.
3. Handling long audio
FFprobe checks that the file contains decodable audio. FFmpeg converts one section at a time to mono, 16 kHz PCM WAV. Each section is at most 30 seconds long, with one second of overlap to reduce words lost at a boundary.
Sections go to Gnani’s REST API sequentially, below its 60-second request limit. Exact repeated word sequences at boundaries are removed before the transcript is assembled. Timestamps refer to chunk boundaries, not word-level alignment. The worker downloads the original to temporary disk and removes temporary files when it finishes or fails.
4. Grounded summaries
Gemini receives the transcript, not the audio or the API key. Its instructions request an overview, key takeaways, topics, and only explicitly stated action items. Responses use a JSON schema and are validated before saving.
Long transcripts are split into bounded text sections, summarized individually, and recursively combined. The transcript is saved before summarization starts. If Gemini fails, the user can read the transcript and retry the summary without paying to transcribe again.
What lives where
PostgreSQL: recording metadata, processing states, transcripts, summary JSON, section checkpoints, and durable job leases.
Object storage: the original audio in a private AWS S3 bucket in both local development and production. Signed URLs grant temporary upload and playback access.
Browser: a signed, HttpOnly workspace cookie and the preferred library view. Audio and transcripts are not stored in localStorage. Separate browser profiles have separate libraries; clearing the cookie loses access to that workspace.
Visible progress and failures
Upload progress comes from actual bytes transferred. During processing, the frontend polls saved state every 2.5 seconds. Transcription progress advances when a section is committed; queued and summarization states explain what is happening without pretending to predict an exact completion time.
Provider timeouts, rate limits, invalid audio, missing configuration, and storage outages surface as actionable messages. Provider calls use bounded retries with backoff. A failed recording retains its original audio and completed transcript sections for an explicit retry.
Tradeoffs and next steps
This development version uses an anonymous, browser-specific workspace so it can be tried without a sign-up flow. The deployment configuration adds HTTPS through Caddy and keeps PostgreSQL on a persistent Docker volume. Further improvements include durable accounts, upload and usage quotas, abuse protection, automated off-host backups, metrics, and an object-deletion outbox. The current PostgreSQL queue is sufficient for a small deployment; higher volumes may justify a dedicated queue and worker autoscaling.
More time would also improve transcript boundaries using silence-aware segmentation, add resumable uploads across browser restarts, compress long-audio playback to a browser-friendly derivative, and provide word-level timestamps where the provider supports them. Currently a cancelled or closed browser upload must be restarted; bucket lifecycle rules clean abandoned multipart uploads after one day. A host crash may leave temporary worker files until the container is replaced.
The project intentionally keeps Next.js, FastAPI, PostgreSQL, object storage, and background jobs separate. Docker Compose runs the frontend, API, worker, and PostgreSQL on one host. Production adds Caddy as the only public entry point on ports 80 and 443. AWS S3, Gnani, and Gemini stay external. One EC2 instance is a single point of failure; a persistent volume is not a backup.