Object / Blob Storage
This page documents storage backends, blob upload routing, and core Docker mount behavior.
Scope of this page
- Focus: object/blob backends, keyspaces, upload/read paths, and storage debugging.
- Not covered here: relational metadata tables and SQL state modeling (see Database).
Storage backends
- Embedded (default): embedded SeaweedFS (
weed mini) blob storage. - External: external S3-compatible object storage.
Metadata database mode (SQLite vs Postgres) is configured separately in Database.
OpenReader currently pins embedded SeaweedFS to 4.18 in CI and Docker builds.
4.19 introduced intermittent InternalError responses on S3 PutObject in our upload flow.
Storage variables are documented in Environment Variables.
Ports
3003: OpenReader app and API routes8333: embedded SeaweedFS S3 endpoint for app/worker storage traffic
The embedded default uses same-origin OpenReader proxy routes, so 8333 does not need to be exposed to browsers.
Upload behavior
OpenReader chooses one browser transport before every transfer with S3_BROWSER_TRANSPORT; it never retries a failed direct request through the app proxy.
proxy(the embedded default): uploads use/api/documents/blob/upload, reads use/api/documents/blob/get, and previews use/api/documents/blob/preview.presigned: browser transfers use signatures generated fromS3_PUBLIC_ENDPOINT. The server and compute worker still useS3_INTERNAL_ENDPOINT.auto: choosesproxyfor embedded SeaweedFS, otherwise requires a public HTTPSS3_PUBLIC_ENDPOINTand choosespresigned.
For a public SeaweedFS/S3 origin, use a dedicated HTTPS hostname such as s3.reader.example, not a /s3 path mount. Preserve the signed path, query string, Host, and signed headers in the reverse proxy. Configure CORS for the OpenReader origin with GET, HEAD, PUT, and OPTIONS, allowing Content-Type and x-amz-server-side-encryption.
Browser Cache Storage
The browser may retain reusable document, preview, and TTS audio responses in the versioned openreader-blobs-v1 Cache Storage cache. This is strictly an evictable performance optimization:
- The server database and object storage remain authoritative.
- Clearing or losing Cache Storage must not change application correctness.
- Cache keys are same-origin synthetic identities and are not fetchable server routes.
- Successful full
200responses may be cached; partial, opaque, redirect-error, and failed responses are not. - Presigned URLs are network sources only and are never used as persistent cache identities.
Synthetic key layouts:
/openreader-cache/documents/{documentId}/{contentVersion}/openreader-cache/previews/{documentId}/{previewVersion}/openreader-cache/audio/{audioKey}/{version}
Explicit audiobook MP3 exports are not persistently cached.
Document previews
- PDF/EPUB previews are generated server-side and stored in object storage under
document_previews_v1. - Preview generation is triggered on upload registration and also backfills on first preview request for older docs.
- Preview artifacts are temporary-cache friendly and can be regenerated from the source document blob.
FS / Volume Mounts
App data mount
- Target:
/app/docstore - Recommended: yes, for persistence
- Purpose: persists SeaweedFS blob data, SQLite metadata DB, migrations, and local runtime temp state
- Mount string:
-v openreader_docstore:/app/docstore
Library source mount (optional)
- Target:
/app/docstore/library - Recommended: optional, use read-only (
:ro) - Purpose: exposes host files as a source for server library import
- Mount string:
-v /path/to/your/library:/app/docstore/library:ro - Details: Server Library Import
Transport topology
The default embedded topology is Browser → OpenReader → http://127.0.0.1:8333 SeaweedFS. For public object storage, use Browser → https://s3.example → S3/SeaweedFS and configure the app and worker with a private S3_INTERNAL_ENDPOINT.
TTS Playback Storage
Worker-owned TTS playback artifacts are stored under dedicated playback keyspaces.
Typical key layout:
${S3_PREFIX}/tts_playback_segments_audio_v1/users/<url-encoded-user-id>/docs/<document-id>/<document-version>/<settings-hash>/<audio-content-hash>.mp3${S3_PREFIX}/tts_playback_segments_v1/users/<user-hash>/docs/<document-id>/<document-version>/<settings-hash>/segments/<ordinal>.json${S3_PREFIX}/tts_playback_plan_v1/...${S3_PREFIX}/tts_playback_v1/...
Notes:
- Docker, Compose,
pnpm dev, andpnpm startautomatically purge the legacytts_segments_v1/,tts_segments_v2/, andaudiobooks_v1/roots during startup. Vercel/custom app-only deployments run the same idempotent cleanup withpnpm migrate-decommissionduring the v5 rollout. - The playback sidecar prefix uses a hashed user id; audio uses the storage user id for content-addressed deduplication.
Account Deletion Cleanup
Account deletion performs best-effort object cleanup:
- Document blobs + preview artifacts
- TTS playback audio and sidecar blobs
If object deletion fails, account deletion still proceeds and orphaned objects may require manual cleanup.
TTS Playback Storage Debug Commands
Use these commands to inspect playback audio objects.
- AWS S3
- Embedded / MinIO / R2 / etc
# List all playback audio objects
aws s3 ls "s3://$S3_BUCKET/$S3_PREFIX/tts_playback_segments_audio_v1/" --recursive
# Filter to one document id (replace <document-id>)
aws s3 ls "s3://$S3_BUCKET/$S3_PREFIX/tts_playback_segments_audio_v1/" --recursive | grep "/docs/<document-id>/"
# List all playback audio objects
aws s3 ls "s3://$S3_BUCKET/$S3_PREFIX/tts_playback_segments_audio_v1/" --recursive --endpoint-url "$S3_INTERNAL_ENDPOINT"
# Filter to one document id (replace <document-id>)
aws s3 ls "s3://$S3_BUCKET/$S3_PREFIX/tts_playback_segments_audio_v1/" --recursive --endpoint-url "$S3_INTERNAL_ENDPOINT" | grep "/docs/<document-id>/"