| last_edited | 2026-10-08 |
|---|---|
| title | Configuration |
| description | Configuration file reference, environment variables, and file locations. |
Default location:
| Platform | Path |
|---|---|
| macOS / Linux | ~/.msgvault/config.toml |
| Windows | C:\Users\<you>\.msgvault\config.toml |
Override the data directory with the MSGVAULT_HOME environment variable or the --home flag (see below).
For a first archive, add only the sections required by your source. Optional provider setup is covered in recommended configuration. The complete example below illustrates the available sections; it is not a required starting configuration.
Starting in v0.20.0, remote deletion remains permanently opt-in. The invoking
CLI may enable it durably with [deletion] remote_enabled = true or for one
command with MSGVAULT_ENABLE_REMOTE_DELETE=1. Both mechanisms are permanent;
there is no planned automatic removal of the guardrail.
Consent belongs to the invoking CLI. When a command uses a remote daemon, the
CLI forwards its effective consent for that operation; the remote daemon's own
[deletion] section is not server policy for a command invoked elsewhere.
Staging, listing, inspecting, and dry-running deletion batches remain ungated.
Use recommended configuration for a guided setup, then return here for exact keys and defaults. Each processing feature has a separate scope and consent contract.
The defaults in this reference apply when a key is absent. They differ from the
values setup providers writes after confirmation: embeddings and people sweeps
start disabled, while setup can enable them and add schedules. Existing explicit
values remain in effect. Provider keys alone do not enable processing.
- Profile automation: tracked people, sweep providers, budgets, and fact resolution.
- Conversation briefs: enrolled people and versioned summaries through the sweep provider.
- External enrichment: Exa/SixtyFour policy setup, exact consent, request limits, and the persistent suppression key.
- Document indexing: extraction, document vectors, and separate query consent.
- Vector search: text, person, and visual indexes.
Identity scoring is disabled by default. An operator starts each batch and
consents to the exact disclosure shown by msgvault person scoring status.
Scoring creates review suggestions and records judgments; it never accepts
matches or links participants.
[people.identity_scoring]
enabled = false
model_id = "jev-1.13.0"
minimum_probability = 0.80
credential_env = "MSGVAULT_JEV_API_KEY"
batch_size = 20
retention_declaration = "provider retention policy accepted by the operator"| Key | Default | Description |
|---|---|---|
enabled |
false |
Enable consented identity scoring. Consent is still required. |
model_id |
jev-1.13.0 |
Fixed provider model identifier. |
minimum_probability |
0.80 |
Probability must be strictly greater than this threshold before local policy can propose acceptance. Accepted values are at least 0.80 and less than 1.00. |
credential_env |
empty | Name of the environment variable holding the provider key. The key value is read from the daemon environment and is never stored in config.toml. Required when enabled. |
batch_size |
20 |
Default and maximum scoring batch size, from 1 through 100. |
retention_declaration |
empty | Operator's exact declaration of the provider retention policy. Required when enabled and included in the consent fingerprint. |
The provider endpoint is fixed at https://api.typesafe.ai/v1/systemone, with
model jev-1.13.0. The disclosure binds that endpoint, model, packet schema,
retention declaration, policy version, and question version. A change to any
of them requires consent to the new fingerprint. status also prints the raw
identity fields and limits covered by the packet schema.
msgvault person scoring revoke <fingerprint> withdraws consent for that
disclosure. It leaves the configuration enabled and retains prior judgments.
Set enabled = false to disable scoring in the configuration. See
identity scoring for the workflow
and the API reference for
endpoints.
People sweeps use one named protocol profile at a time. A profile records the exact endpoint, model, wire protocol, negotiated output mode, privacy posture, and source scope. Built-in OpenAI, OpenRouter, and Venice presets bind the protocol, endpoint, and authentication scheme; you still choose the model and privacy policy. Msgvault never changes the active profile or switches providers automatically.
[people.sweep]
enabled = true
provider = "glm"
[people.sweep.providers.glm]
protocol = "openai_chat"
endpoint = "https://api.z.ai/api/paas/v4"
model = "glm-5.3"
auth = "bearer"
credential = "env"
credential_env = "ZAI_API_KEY"
output_mode = "prompt_json"
token_limit_parameter = "max_tokens"
reasoning_effort = "max"
request_timeout = "1m"
retention_posture = "provider-declared"
training_posture = "provider-declared"
allowed_sources = ["conversation_text", "meeting_text"]
source_since = "2026-01-01"
allow_sensitive = trueallow_sensitive = true is required for real sweeps: any packet that
carries seed or context evidence is marked sensitive, so a profile set to
false fails on every real sweep. true permits sending that verbatim
archive text to the selected provider; false leaves only the synthetic
capability check, which sends no archive text.
Usable protocols are openai_chat, openai_responses,
anthropic_messages, and google_generate_content. A fifth protocol,
codex_app_server, is defined but release-gated and cannot run yet; see the
Codex app-server profiles section below. Onboarding negotiates and saves
native_json_schema, json_object, or prompt_json. OpenAI Chat profiles
also save either max_completion_tokens or max_tokens; the other protocols
use their defined token-limit field.
Other providers use explicit protocol profiles; OpenRouter, Venice, and OpenAI also have built-in presets:
| Example profile | Protocol | Typical profile choice |
|---|---|---|
| GLM 5.3 | openai_chat |
Z.AI API base, glm-5.3, often max_tokens |
| Kimi K3 | openai_chat or anthropic_messages |
Choose the exact API surface the account exposes |
| OpenRouter | openai_chat |
OpenRouter API base and one explicit routed model ID |
| Venice | openai_chat |
Venice API base and one explicit model ID |
| open-agent-api | openai_chat |
The gateway's loopback API base and exposed model ID |
| Gemini | google_generate_content |
Google API base and one Gemini model ID |
| Anthropic | anthropic_messages |
Anthropic API base and one Claude model ID |
| OpenAI Responses | openai_responses |
OpenAI API base and one Responses model ID |
Confirm current endpoints, model identifiers, privacy terms, and subscription rules with the selected operator before saving a profile. OpenRouter and Venice may route a request to another upstream operator, so the profile's retention and training declarations must cover that full path. Logged-in or subscription-backed endpoints, including local gateways, must be used within their provider terms.
Credentials are not stored in this TOML. credential = "stored" keeps a
profile-specific secret in tokens/provider-credentials.json on every
platform, sent only to the endpoint origin of the profile it was saved for.
If you edit a profile's endpoint to a different origin, its stored key stops
working; remove the profile and add it again with the new endpoint.
Keys that an older release stored under tokens/people-providers/ on Linux or
macOS move into that file the first time msgvault uses the profile.
credential = "env" stores only the selected environment-variable name.
Environment-variable names are host-only settings: configure them through the
CLI or TOML, not the Web UI.
credential = "none" is restricted to credentialless local or Codex paths.
Changing a credential value does not change the profile fingerprint, but
changing its source or reference does.
Only msgvault person provider add may contact models.dev, and only when a
transport field (--protocol, --endpoint, --model, --auth) is missing
or --accept-catalog-prices is set; it sends no archive data or provider
credential. --custom skips that catalog entirely; the required synthetic
check still contacts the endpoint selected in the profile. The catalog is never used
by scheduled or manual sweeps. A catalog suggestion also never chooses where a
credential is sent: onboarding pairs a credential only with an endpoint you
passed explicitly via --endpoint or with the first-party API hosts compiled
into msgvault, so a compromised catalog cannot redirect your key. A successful check does not grant consent:
msgvault person provider consent <name> --yes is a separate explicit step.
Live credential checks are optional developer or operator verification and are
never CI requirements.
Enable model-assisted profile maintenance only after configuring, checking, and consenting to a provider. A person must also be tracked before the sweep maintains their facts. Conversation briefs use this same provider and schedule, with separate enrollment and interval controls.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Run the scheduled people sweep with the selected provider. |
provider |
default |
Name of a table under [people.sweep.providers]. The initial profile has an OpenAI endpoint but no model; it is not a usable, consented provider. Setup can explicitly select openai, openrouter, or venice, or configure local ollama. |
schedule |
15 2 * * * |
Daily at 02:15 in the daemon's time zone. An omitted or empty value receives this default; use enabled = false to disable the sweep. |
work_batch_size |
25 |
Tracked people considered in one worker batch. |
historical_message_cap |
2000 |
Maximum archived messages considered when finding context for each profile field. |
context_per_target |
8 |
Maximum context items selected for each profile field. |
evidence_max_bytes |
131072 |
Byte limit for an evidence packet. |
evidence_max_items |
200 |
Item limit for an evidence packet. |
backstop_interval |
24h |
Interval before checking tracked people for changes missed by incremental work. |
Hosted people inference requires an explicit setup providers --provider <openai|openrouter|venice> --model <model> choice, a credential source, and
explicit retention, training, and sensitive-content decisions. An
OPENAI_API_KEY alone configures only eligible embedding lanes. With no OpenAI
key or explicit provider choice, setup can offer the configured loopback Ollama
chat model when --allow-sensitive is supplied. An enabled sweep is preserved;
--provider then fails with instructions to use person provider add and
person provider use. See setup flags.
Request and token limits apply to sweeps and briefs. Setup keeps these defaults; it does not obtain provider prices or set a monetary limit.
| Scope | Request key and default | Input-token key and default | Output-token key and default |
|---|---|---|---|
| One person | max_requests_per_person = 4 |
max_input_tokens_per_person = 200000 |
max_output_tokens_per_person = 16000 |
| One run | max_requests_per_run = 100 |
max_input_tokens_per_run = 1000000 |
max_output_tokens_per_run = 160000 |
| One day | max_requests_per_day = 500 |
max_input_tokens_per_day = 5000000 |
max_output_tokens_per_day = 800000 |
max_estimated_cost_microusd_per_run and max_estimated_cost_microusd_per_day
both default to 0, which disables those cost limits. To use either, supply
positive input_cost_microusd_per_million_tokens and
output_cost_microusd_per_million_tokens; both price assumptions also default
to 0. Values are integer millionths of a US dollar. These are local estimates
from the prices you provide, not a provider billing limit.
Control how often enrolled people receive a "Last time we talked" brief and
how much text each generation uses. Briefs use the selected sweep provider and
share its budgets. The profile must permit sensitive content and include
conversation_text in allowed_sources.
Only supported chat and text-message sources supply brief evidence; email, meeting transcripts, documents, and your own replies are excluded. See the brief guide for supported sources and enrollment instructions.
Generation is enabled here by default, but each person must be enrolled separately. These settings apply to everyone; there are no per-person overrides. Restart the daemon after changing them.
[people.sweep.brief]
enabled = true
min_interval = "168h"
pre_call_window = "72h"
max_items = 40
max_bytes = 65536
overlap_items = 8
max_output_tokens = 2048
max_rendered_runes = 560| Key | Default | Description |
|---|---|---|
enabled |
true |
Generate briefs for enrolled people when the sweep runs. false stops generation for everyone without removing enrollments. |
min_interval |
168h |
Minimum age of the current version before a scheduled run regenerates it. A rejected current version counts as no brief and is replaced on the next eligible run. msgvault person brief generate bypasses this. Must be positive. |
pre_call_window |
72h |
Regenerate this far ahead of a due contact cadence, so the brief is current before you reach out. Must be positive. |
max_items |
40 |
Maximum archive items admitted to one brief window. Must be positive. |
max_bytes |
65536 |
Maximum packet size for one brief window. Must be positive. |
overlap_items |
8 |
Items already covered by the previous brief's window that may be re-admitted so a continued thread is recognizable. Must not be negative or exceed max_items. |
max_output_tokens |
2048 |
Output cap for the brief call. Must be positive and must not exceed [people.sweep.budgets] max_output_tokens_per_person, because a brief is one more call against the same per-person ceiling. |
max_rendered_runes |
560 |
Maximum length of the rendered paragraph, in Unicode runes. Msgvault enforces it after the model answers by dropping structured items from the tail: uncertainties first, then follow-ups, then highlights, never the last-interaction sentence and never mid-sentence. Every dropped item is counted in the version's dropped_item_count, and the stored structure and evidence pointers are the trimmed ones. Must be at least 240, which is the interaction summary's own maximum length. |
An invalid value fails configuration validation with the offending key named, rather than being clamped.
The codex_app_server protocol is not usable in this release. Its transport
stays unavailable until the executable isolation gate releases a verified
build, and until then every Codex operation fails closed with
codex app-server isolation is not released. The profile shape is documented
here so the configuration is ready when the gate ships. Codex sign-in and
model routes return HTTP 503 before changing credentials or consent. The
terminal-only person provider enroll-codex
command creates a new profile through the daemon; host-side person provider login reauthenticates the selected existing Codex profile. Both remain gated.
The Linux launcher disables Codex's local execution environment and shell tools. The app server can manage its staged OAuth credential, but its command and filesystem interfaces cannot access it. Command-execution requests or events abort inference. Enabling a release still requires a real authenticated structured-inference check through this launcher.
codex_app_server profiles are also the one protocol person provider add
cannot create: generic onboarding negotiates HTTP capabilities through an
endpoint, while codex_app_server has no endpoint to negotiate against and
runs through an attested local Codex executable instead. The following is
a reference shape for that gated implementation, not a working setup procedure:
[people.sweep.providers.codex]
protocol = "codex_app_server"
model = "gpt-5.3-codex"
auth = "none"
credential = "none"
reasoning_effort = "medium"
retention_posture = "zero_retention"
training_posture = "no_training"
allowed_sources = ["conversation_text"]
source_since = "2026-01-01"
allow_sensitive = trueendpoint is not allowed, and auth and credential must both be set to
"none": the transport is the local Codex app server, authenticated by its
own ChatGPT login. Sensitivity policy is protocol-agnostic: real Codex sweeps
send the same seed and context packets, so allow_sensitive = true is
required here as well. Login, model discovery, checks, and consent cannot make
this profile usable while the release gate is closed. Use one of the available
HTTP protocols for current profile automation.
TOML treats backslashes inside double-quoted strings as escape characters. On Windows, this means native paths like "C:\Users\you\..." will cause a parse error.
Use one of these formats instead:
# Forward slashes (recommended)
client_secrets = "C:/Users/you/Downloads/client_secret.json"
# Single-quoted string (backslashes are literal)
client_secrets = 'C:\Users\you\Downloads\client_secret.json'msgvault setup providers writes recommended values for the [vector],
[attachments.documents], and [people.sweep] sections from the API keys in
your environment, and msgvault setup status reports every lane with its
provider, model, consent state, and next step. The values it chooses are
listed in Recommended Configuration.
| Key | Default | Description |
|---|---|---|
data_dir |
~/.msgvault |
Base directory for all data |
export_dir |
{data_dir}/exports |
Directory for attachment ZIPs, downloads, and files opened from the TUI |
database_url |
{data_dir}/msgvault.db |
SQLite database path or PostgreSQL DSN |
loose_attachments |
false |
Keep attachments as loose files and reject pack/repack commands instead of creating immutable packs |
Attachments and OAuth tokens are stored in subdirectories of data_dir (attachments/ and tokens/ respectively). These paths are not independently configurable.
Setting loose_attachments = true prevents new pack files but does not
convert existing packs. Stop the daemon and run msgvault unpack-attachments
once to materialize their contents as loose files. Backup restore also restores
attachments loose while this setting is enabled.
Hosted extraction and local full-text indexing for standalone document attachments. It is disabled by default. Enabling it does not grant consent or send data: an operator must generate an authenticated capability manifest and record consent for the exact effective policy before a build can upload a document.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Allow explicit document extraction commands |
provider |
mistral |
Pinned extraction provider |
region |
eu |
Pinned provider region and EU endpoint |
api_key_env |
MISTRAL_API_KEY |
Environment variable containing the provider key |
model |
mistral-ocr-4-0 |
Pinned OCR model |
retention_posture |
unknown |
Confirmed provider posture: standard or zdr |
training_posture |
unknown |
Confirmed provider posture: default-opt-out or opted-out |
max_file_bytes |
52428800 |
Maximum original document size (50 MiB) |
max_pages_per_document |
500 |
Maximum provider units for one document |
max_response_bytes |
67108864 |
Maximum provider response size (64 MiB) |
max_normalized_chars |
25000000 |
Maximum locally retained normalized characters |
max_spool_bytes |
536870912 |
Maximum private staging-directory usage (512 MiB) |
min_free_space_bytes |
1073741824 |
Free space preserved before staging (1 GiB) |
request_timeout |
5m |
Timeout for each provider request attempt |
max_retries |
3 |
Maximum transient retries |
max_pages_per_run |
10000 |
Conservative provider-unit budget for one run |
max_estimated_cost_usd_per_run |
50 |
Cost-planning ceiling for one run |
estimated_cost_usd_per_1000_units |
0 |
Operator-supplied current price assumption; zero disables cost calculation |
pricing_assumption_on |
— | Date for the price assumption, in YYYY-MM-DD form |
| Key | Default | Description |
|---|---|---|
enabled |
false |
Convert standalone text/csv attachments locally to PDF before the authorized PDF upload; disabled conversion leaves raw CSV outside the provider-authorized scope |
CSV conversion uses Docbank's default record, cell, cell byte, and PDF limits, tightened by the configured original file, response, and page ceilings. The generated PDF is transient. The archive keeps the original CSV hash and MIME type plus the conversion receipt and page, record, and cell provenance. The conversion declaration participates in the exact consent fingerprint only when enabled.
Provider uploads are manual-only: msgvault serve does not schedule document
extraction. Each documents build or documents resume receives its capability
manifest explicitly and displays its upload and cost preflight before requiring
--yes. When document indexing is enabled, the daemon's weekly reconciliation
and local derivative cleanup remain automatic and make no provider requests.
[attachments.documents.scope] accepts these fields:
| Field | Default | Meaning |
|---|---|---|
message_types |
[] |
Include all supported message sources, or restrict extraction to the listed types |
include_inline |
false |
Also include inline attachments with an authorized document media type and authoritative role provenance |
Inline scope support is available in v0.21.0. Some mail clients mark
ordinary document attachments as inline. Set include_inline = true to include
them; other roles remain excluded. This changes the consent fingerprint. Run
msgvault documents consent-mistral --capabilities <manifest> --yes again before
building. Selecting a standalone-only profile stops inline search results from
serving, including results extracted under an earlier profile.
The first release requires
[attachments.documents.index].lexical = true and store_chunk_text = true.
Hosted document embeddings are not enabled by this configuration.
Enable document vectors separately with
[attachments.documents.index.embeddings] enabled = true, an enabled text
embedding provider, and distinct consents for document text and query text.
They use [vector.embed.schedule] for automatic embedding of already extracted
document chunks. That schedule never extracts new attachments. Setup enables
this subtable when both document extraction and a supported text provider are
selected, but leaves the two consent steps to you.
See Document Attachment Indexing for the complete probe, consent, build, and recovery flow.
| Key | Default | Description |
|---|---|---|
client_secrets |
— | Path to Google OAuth client_secret.json for browser OAuth flows |
service_account_key |
— | Path to a Google service account key JSON for Workspace domain-wide delegation |
Named OAuth apps for Google Workspace organizations that require their own OAuth credentials. Each entry can define a separate browser OAuth client_secret.json, service account key, or both. Use --oauth-app <name> with add-account to bind an account to a named app.
| Key | Default | Description |
|---|---|---|
client_secrets |
— | Path to the org's client_secret.json |
service_account_key |
— | Path to the org's Google service account key JSON |
See OAuth Setup: Google Workspace Accounts for when and why you need named apps.
Discord's --oauth-app value is only a protected bot-token binding label. It
is not resolved from this section and does not require an [oauth.apps] entry.
export-messages uses the same daemon and database configuration as other
archive commands. It does not load provider credentials or make provider API
calls. The older export-discord compatibility command has the same read-only
provider behavior.
When service_account_key is configured, msgvault add-account <email> validates the delegated Gmail profile and registers the account without storing a per-user refresh token. The service account key file must be owner-only on Unix-like systems, for example chmod 600 /path/to/service-account.json.
[carddav] is the default connection. Add named tables for other accounts;
each uses the keys below and has its own credential binding, discovery,
retry state and schedule. Names use 1–64 lowercase ASCII letters, digits,
underscores or hyphens, starting with a letter. default is reserved for
[carddav].
Connect through the CardDAV account workflow so the
daemon validates discovery before saving these settings. The same base_url
and username cannot belong to two connections, including disabled connections.
For Google, account email is case insensitive and the OAuth app does not create
a separate CardDAV account. Config edits and account saves reject duplicates.
Removing a config table retains the account's archive data; see
recovering an orphaned connection.
| Key | Default | Description |
|---|---|---|
provider |
"" |
Empty for a password-based server, google for Google Contacts, or microsoft for Microsoft 365 and Outlook.com contacts |
oauth_app |
"" |
Named Google OAuth app; empty selects [oauth] |
base_url |
"" |
CardDAV discovery URL; Google and Microsoft setup supply their canonical URL |
username |
"" |
Server username, or Google or Microsoft account email |
schedule |
"" |
Cron schedule; empty disables scheduled sync |
enabled |
false |
Enable the configured connection |
trusted_origin |
"" |
Exact HTTPS origin approved for private access, including its port; a trailing / is accepted. Applies only when it matches the account URL's origin. |
trusted_addresses |
[] |
Private IP addresses to dial for trusted_origin, without DNS. Accepts 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10, and fc00::/7; rejects duplicates, IPv6 zones, loopback, and link-local addresses. |
For example, a second connection uses:
[carddav_connections.work]
base_url = "https://contacts.example.com/dav/"
username = "you@example.com"
enabled = true
schedule = "0 */6 * * *"Use add-carddav --connection work or Add connection in Settings to save
its credential and discover books. Passwords are rejected in these TOML tables.
The default binding remains tokens/carddav.json; named bindings use
tokens/carddav-connections/<name>/carddav.json. Google authorizations are
shared by account email and OAuth app, separately from these connection bindings.
Scheduler job names are carddav for default and carddav:<name> otherwise.
Set both trusted-destination keys together. See the private-server setup for an example, restart requirements, and behavior when the origin does not match.
Passwords and Google tokens stay in the configured token directory, outside
config.toml. See Google Contacts setup
for browser and terminal authorization.
Configuration for Microsoft 365 / Outlook.com OAuth and Microsoft Teams Graph
sync. Required only if you use add-o365, add-teams, or sync-teams.
| Key | Default | Description |
|---|---|---|
client_id |
— | Azure AD Application (client) ID (required) |
redirect_uri |
http://localhost:8089/callback/microsoft |
OAuth redirect URI registered in the Azure AD app |
tenant_id |
common |
Azure AD tenant ID; common allows both personal and org accounts |
See OAuth Setup: Microsoft 365 for app registration steps. Teams uses the same client_id but requests Microsoft Graph scopes and stores tokens under tokens/teams_<email>.json; Outlook/Hotmail IMAP OAuth uses tokens/microsoft_<email>.json.
Optional source-scoped Fastmail JMAP identity inventory. This does not replace IMAP ingestion credentials: add and sync the mailbox normally, then use the API token only to discover masked and send-as addresses that belong to that source.
| Key | Default | Description |
|---|---|---|
source_id |
— | Positive numeric archive source ID; mutually exclusive with account |
account |
— | Unambiguous source identifier or display name; mutually exclusive with source_id |
api_token |
(required) | Fastmail API token used for the JMAP identity inventory |
auto_confirm_identities |
false |
Refresh and apply strong provider identity evidence after successful mailbox syncs |
Exactly one source selector is required. Prefer source_id when two sources
share an identifier or display name. With automatic confirmation disabled,
msgvault identity discover --source-id <id> --provider fetches the inventory
for an explicit preview; add --apply only after reviewing it. See People,
Profiles, and Source Identities.
Provider-wide Discord import settings and optional message-container filters.
Register guilds and store their bot credential first with msgvault add-discord; tokens and binding labels do not belong in config.toml.
| Key | Default | Description |
|---|---|---|
max_media_bytes |
52428800 (50 MiB) |
Maximum size of one Discord attachment downloaded during sync or backfill |
max_media_mb |
— | Same cap in MiB; when set it takes precedence over max_media_bytes |
media |
true |
Download attachment bytes at all |
media_scope |
all |
Which conversations collect media: all, direct (direct and group chats only), or none |
media_max_participants |
20 |
Skip media from conversations with more participants than this; 0 disables the cap. See Media policy |
edit_rescan_window |
168h (seven days) |
Trailing per-channel/thread window refreshed for edits, deletions, and reaction summaries |
Use an exact guild ID for a per-guild filter block. The same block also takes
the per-account media overrides (media, max_media_mb) that the other chat
providers put under accounts_config:
[discord.guilds."123456789012345678"]
include = ["456789012345678901"]
exclude = ["567890123456789012"]
# media = false
# max_media_mb = 25An empty include means every accessible text or announcement channel, thread,
and forum post. Top-level channels match directly. A child inherits its
parent's state unless its own ID appears explicitly. An explicit child include
can override an excluded parent; an explicit child exclude can override an
included parent. exclude wins when the same ID is in both lists. See
Discord.
[beeper], [slack], [discord], and [teams] share one attachment policy
vocabulary. It decides which chat media is downloaded during sync and backfill;
message text is always archived.
| Key | Default | Description |
|---|---|---|
media |
true |
Download attachment bytes. false archives messages without their media and records a policy_scope skip marker |
media_scope |
all |
all collects from every conversation; direct collects only from direct and group chats (not channels, rooms, or guild channels); none collects nothing |
media_max_participants |
20 |
Skip media from conversations with more participants than this. Omitting the key applies the default; an explicit 0 removes the cap |
max_media_mb |
250 (Discord 50) |
Per-attachment size cap in MiB. Sized for long voice notes, screen recordings, and phone video from direct chats now that the participant cap keeps large-room volume out |
accounts_config |
— | Per-account overrides of media and max_media_mb, keyed by Beeper accountID, Slack team ID, or Teams account email. Discord uses [discord.guilds."<id>"] instead |
The participant cap exists because most attachment bytes in a real chat
archive come from large rooms whose forwarded videos nobody wants kept. Direct
chats and small groups keep their photos, voice notes, and files. A skipped
occurrence is recorded with a typed marker (participant_threshold,
policy_scope, account_policy, or size_cap) that distinguishes a
deliberate skip from a failed download, so the backfill-*-media commands do
not retry it unless the policy changes.
[beeper]
media_scope = "all"
media_max_participants = 20
max_media_mb = 250
# Keep everything from one account regardless of room size or size cap.
[beeper.accounts_config.signal]
media = true
max_media_mb = 500
# Never download from another account.
[beeper.accounts_config.telegram]
media = falsePolicy changes apply to future downloads. Media already stored under an
earlier policy stays until you run msgvault purge-excluded-media, which
removes attachment bytes the current policy would no longer collect.
Structured file logging. Disabled by default. Enable it to get persistent, machine-readable logs for troubleshooting. Every CLI invocation writes a unique run_id on every log line so you can trace a single run across shared daily log files.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on persistent file logging. Setting dir also enables it implicitly. |
dir |
<data_dir>/logs |
Directory for log files |
level |
info |
Log level: debug, info, warn, error |
sql_trace |
false |
Log every SQL query at info level (verbose, for debugging) |
sql_slow_ms |
100 |
Threshold in ms above which SQL queries are logged at warn level. 0 uses the built-in default (100 ms). |
Log files are named msgvault-YYYY-MM-DD.log (UTC date), written as newline-delimited JSON. When a daily log exceeds 50 MiB it rotates to .log.1, .log.2, etc. (up to 5 rotated files).
When SQL logging is enabled, slow/error entries include query arguments and streaming query durations, which makes it easier to diagnose expensive reads without enabling full trace output.
Use msgvault logs to view and tail log files from the selected local or remote daemon. See CLI Reference: logs.
| Key | Default | Description |
|---|---|---|
rate_limit_qps |
5 |
Scales Gmail's local quota-unit refill rate: 5 allows 250 units/second; 3 allows 150. Gmail values above 5 are capped at 5. Also sets Microsoft Teams Graph requests/second, without that cap, so lowering it slows Teams imports too. Reduce it if Gmail reports quota errors; Google's project quotas can be lower than this local budget. |
archive_remote_images |
false |
Download remote email images during Gmail/IMAP sync and EML, EMLX, MBOX, and PST imports |
trusted_imap_sent_mailboxes |
{} |
Per-IMAP-account Sent-folder names (keyed by the ACCOUNT identifier from msgvault list-accounts) that enable edited-copy snapshot refresh for servers without advertised special-use roles |
Remote image archiving is off by default. Enabling it contacts sender-controlled servers and can activate tracking pixels or disclose the archive server's IP address. Restart the daemon after changing the setting. It applies to new ingestion; existing mail needs an explicit backfill.
See remote email images for the opt-in workflow, supported formats, download limits, and effect on attachment counts.
trusted_imap_sent_mailboxes names each IMAP account's Sent folder for
servers — such as some Exchange/Outlook accounts — whose (possibly localized)
Sent folder advertises no RFC 6154 \Sent special-use role. A survivor copy in
that unadvertised folder never replaces the archived snapshot by default:
the sync adopts the surviving location but keeps the previously archived
body, raw MIME, recipients, and attachments, because a sender can forge any
RFC822 Message-ID. Placement the server does advertise — an unambiguous
\Sent or \Drafts role — remains trusted automatically regardless of
this setting. Naming a mailbox here states that it is that account's Sent
folder and holds only mail the account itself authored, which re-enables
refreshing the archived snapshot from an edited copy found there.
The mapping is keyed by the exact source identifier — copy the ACCOUNT
value printed by msgvault list-accounts (for example
imaps://user@example.com@imap.example.com:993), not the email or display
name. Trust never crosses accounts: a same-named mailbox in another synced
account stays untrusted, and an account with no entry has no explicit trust.
[sync]
trusted_imap_sent_mailboxes = { "imaps://user@example.com@imap.example.com:993" = ["Gesendete Elemente"] }This is an explicit trust assumption, not evidence: filters or IMAP rules
that file received mail into the listed mailbox would let that mail replace
an archived snapshot under the same Message-ID. Advertised unambiguous
\Sent and \Drafts placement is trusted automatically, and a mailbox
that carries \All, \Junk, or \Trash roles, or INBOX, is never
trusted — not even when listed here explicitly. A configured name the server
itself advertises as \Drafts keeps its Drafts meaning: explicit
configuration cannot turn a Drafts folder into the account's Sent folder.
Settings for the Web UI and API server started by msgvault serve. The same HTTP server is used by remote CLI access and by the local background daemon for archive-access CLI commands. The api_key setting is also reused for inbound bearer authentication when msgvault mcp --http starts a separate Streamable HTTP listener; that listener's address comes from the --http flag. See Web UI & API Server for API endpoint documentation and MCP Server for MCP client setup, or fetch /openapi.json from a running server for the generated OpenAPI contract.
| Key | Default | Description |
|---|---|---|
api_port |
0 (auto-select) |
Port the server listens on; 0 picks an open port at startup and clients discover it automatically. Set a fixed port for remote/NAS deployments. |
bind_addr |
127.0.0.1 |
Bind address, or iface:NAME to bind an address on a named interface |
api_key |
— | API key for daemon/API authentication and bearer authentication on msgvault mcp --http |
api_key_file |
— | Owner-only file holding the API key |
api_key_env |
— | Name of an environment variable holding the API key |
agent_access |
false |
Enable restricted agent grants; requires an effective API key. Read at daemon startup only; a config.toml edit takes effect only after a restart. |
allow_insecure |
false |
Allow non-loopback binding without api_key |
cors_origins |
[] |
Allowed CORS origins |
cors_credentials |
false |
Allow credentials in CORS requests |
cors_max_age |
0 |
CORS preflight cache duration in seconds |
trusted_proxies |
[] |
IP addresses or CIDRs allowed to supply forwarded HTTPS/host headers |
daemon_idle_timeout |
20m |
Idle timeout for lifecycle-managed background daemons; set to "0s" to disable |
daemon_auto_restart |
newer |
Local daemon restart policy when the CLI finds a different daemon binary version: newer, never, or always |
daemon_auto_start |
true |
Let CLI, TUI, and MCP commands start a local background daemon when none is running; set false when a supervisor runs msgvault serve |
remote_clients |
[] |
Read-only API keys for remote CLIs; requires an effective server API key. See below. |
Each [[server.remote_clients]] entry gives one remote CLI its own read-only key. See read-only remote clients for what those keys can do and their limits.
| Key | Default | Description |
|---|---|---|
client_id |
— | Unique name, used in server logs |
api_key_file |
— | Owner-only file holding the client's key; relative paths resolve like the server's api_key_file. Must differ from the server key and other clients' keys. |
collections_write |
false |
Also allow creating, editing, and deleting collections |
msgvault serve reads these files after it settles the server key, whether that key comes from api_key, a file, an environment variable, or is created at startup. It fails to start if a file is missing or empty or a key repeats another. CLI commands that start a local daemon also trigger these checks. Restart to add, remove, or rotate a key.
daemon_idle_timeout applies only to background daemons started by msgvault daemon start or auto-started by a CLI command. Foreground msgvault serve keeps running until stopped. MSGVAULT_DAEMON_IDLE_TIMEOUT overrides the configured value for lifecycle-managed background daemons.
On unreleased main, flags and environment variables can configure the server
without config.toml. serve --bind and serve --port take priority over
environment variables, then TOML, then defaults. iface:NAME resolves at
startup, preferring a usable IPv4 address and otherwise IPv6. An unknown, down,
or unaddressed interface fails before opening a listener. Startup logs report
the bound address and whether the bind came from a flag, environment, a config
file, or the default.
Config edits validate the saved settings without resolving network interfaces or reading server credentials. These resources must be available when the server starts; an unavailable interface or key does not block unrelated edits.
Credentials use api_key, then api_key_file, then api_key_env. A selected
file or named variable that is missing, empty, or unsafe fails without trying
another source. Files must be regular, owned by the process user, at most
64 KiB, and readable only by that user (0400 or 0600 on Unix). Symlinks are
rejected. Windows files require an owner-only ACL. File reads trim surrounding
whitespace and leave mounted permissions unchanged. Relative secret paths in
TOML or environment variables resolve beside config.toml, including before
the default file exists. With --config, they resolve beside the selected file.
Saving configuration preserves the original credential path strings, so relative
paths continue to work when the configuration directory moves.
Container secret mounts must meet these same rules. Docker Swarm can set the
secret's uid to the process user and its mode to 0400; see
Swarm secret options.
The default root-owned 0444 mount is rejected. With Docker Compose file
mounts, set the host file's ownership and mode before mounting it.
Kubernetes Secret volumes use symlinks and are not accepted directly. Use a
Secret-backed environment variable, or provide a regular owner-only file.
These restrictions apply to api_key_file, MCP's --http-token-file, and
credentials set --from-file. credentials set --stdin reads a stream and
can import a readable mounted secret through shell input redirection.
When a secure non-loopback server has no configured credential source, it
creates <data_dir>/tokens/server-api-key with an unpredictable key and
owner-only permissions. Persist data_dir to retain the key across restarts.
It reuses that key on later starts, including loopback-only starts. The Web UI
then requires login on 127.0.0.1 too; use the key from that file. An invalid
existing key fails instead of being replaced. Local CLI clients discover the
persisted key, including when they start the daemon themselves.
allow_insecure = true skips this default key and retains the explicit
unauthenticated mode; explicitly configured keys are still enforced. Startup
logs name the credential file without printing its contents.
Run local CLI commands as the daemon's operating-system user and provide the
same selected credential sources. A different user, including root through
sudo or docker exec, fails the key file's ownership check. For a container,
use docker exec --user <daemon-uid> .... If api_key_env names a variable
provided only to the supervised daemon, also provide it to the CLI process;
the CLI does not inherit the daemon's environment.
daemon_auto_restart = "newer" replaces an older compatible local daemon with the current CLI binary. Use "never" when another supervisor owns the daemon lifecycle, or "always" to restart whenever the recorded daemon version differs. Remote servers are never auto-restarted by a CLI client.
daemon_auto_start = false is for installs where a supervisor such as launchd, systemd, or Docker runs msgvault serve. Local archive commands then use the daemon that is already running, wait for one that is still starting, and otherwise fail with an error instead of starting their own. They also never replace a running daemon, whatever daemon_auto_restart says, because the supervisor owns restarts. msgvault daemon start, msgvault daemon restart, and the restart after msgvault update still start a daemon when you run them. Commands routed to [remote].url are unaffected.
Browser sessions are additive to API-key authentication. Existing CLI and
programmatic clients continue to send the configured key. For remote browser
access, terminate TLS at a reverse proxy and list that proxy—not arbitrary
clients—in trusted_proxies. See Web UI for the complete security
model and the plain-HTTP warning.
For MCP Streamable HTTP, send the effective [server] key, or the key selected
by --http-token-file or --http-token-env, as Authorization: Bearer <key>
on every /mcp request. This inbound credential is independent of
[remote].api_key, which authenticates msgvault mcp when it connects to a
remote daemon.
Defaults for the daemon-served browser application. These values can also be
changed from Settings; config.toml remains authoritative.
| Key | Default | Description |
|---|---|---|
default_search_mode |
full_text |
Initial mode: full_text, semantic, or hybrid |
theme |
system |
Color theme: system, light, or dark |
density |
compact |
Table density: compact or comfortable |
Browser-managed settings are validated and written with optimistic concurrency.
Only the [web] keys apply right away; every other config.toml category takes
effect after the daemon restarts, and the Settings page says so once per
category. Two things saved from the Settings page are not config.toml rows and
apply right away: person-enrichment provider API keys, and the CardDAV account,
which has its own save action. A "Saved. Restart the daemon" banner means the
file is saved but the running daemon still has its old value. Changing server.api_key requires a confirmation and takes
effect only after restart, which also invalidates browser sessions.
Optional provider-neutral integration for message-to-task links. Person agendas use the separate Kata connection below.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Enable discovery and capability checks |
endpoint |
— | Explicit loopback HTTP, Unix socket, or HTTPS endpoint; empty requests secure local discovery |
api_key |
— | Server-side credential; the browser sees only a masked hint of it, never the key |
default_project |
msgvault |
Fixed project used for create/link/search operations |
Remote plaintext HTTP is rejected. An endpoint is usable only when it supports the required idempotency and compare-and-swap capabilities; the UI distinguishes disabled, authentication required, incompatible, partial, stale, unavailable, and ready states.
Optional live person agendas backed by Kata, and the connection that files Kata issues from archive evidence. Tasks stay in Kata; msgvault shows their current state when you open a person's agenda. This integration is built against Kata v0.18.0 and requires Kata API schema version 0.21.0 or later.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Enable Kata person agendas and evidence issues |
endpoint |
— | Required when enabled: an explicit HTTPS URL, loopback HTTP URL, or Unix socket URL |
api_key |
— | Bearer credential sent by the daemon to Kata; Settings returns only its configured state and a masked hint |
default_project |
msgvault |
Existing active Kata project used for person agendas and evidence issues |
Create the project in Kata, then configure its endpoint and credential on the machine running the msgvault daemon:
[integrations.kata]
enabled = true
endpoint = "https://kata.example.com"
api_key = "replace-with-your-kata-api-key"
default_project = "msgvault"Restart the msgvault daemon after saving. These settings are also editable in
Settings and take effect after restart. Changing the endpoint to a different
origin in Settings clears its saved key unless you provide a replacement key
in the same save. Remote plaintext HTTP is rejected. Kata does not use local
endpoint discovery, and [integrations.tasks] does not configure person agendas.
Each Kata task can belong to one person. The scalar metadata value
msgvault.person is the person's canonical vCard UID, not their numeric
msgvault person ID. msgvault.list names its list and defaults to agenda.
Reads and unlink operations also recognize the person's UID aliases after a
merge.
Agendas show open tasks only, with at most 100 returned items. A truncated
response means more remain; open Kata to see them. The transport also limits
each response to 1 MiB and reports an error when it exceeds that limit.
Create, link, move between lists, and unlink tasks through msgvault; edit task
content or priority, complete tasks, and reopen them in Kata. See
person agenda commands.
Settings for daemon-side aggregate query behavior. The Web UI, TUI, MCP server, and aggregate list commands use these settings through the local daemon or a configured remote server.
| Key | Default | Description |
|---|---|---|
engine |
auto |
Aggregate engine: auto starts with live SQL and switches to DuckDB after cache maintenance succeeds; sql always uses live SQL; duckdb requires a usable Parquet cache |
auto_build_cache |
true |
Refresh a stale or missing Parquet cache automatically at startup, after scheduled or manual syncs, and when a query finds it due; false skips automatic builds. An explicit query --fresh or sync --build-cache can still request one |
min_rebuild_interval |
0s |
Minimum age of a usable cache before a sync, query, or daemon restart may queue an automatic rebuild. Queries serve the committed snapshot during the interval |
builder_memory_limit |
2GB |
DuckDB buffer-manager budget for cache builds, such as 4GB or 512MiB; total process memory can exceed it |
builder_threads |
min(CPUs, 2) | DuckDB threads for cache builds; zero keeps the default |
builder_temp_limit |
32GB |
Maximum spill-to-disk size for cache builds |
query_memory_limit |
512MB |
DuckDB memory limit for daemon aggregate queries; raise it on a large archive |
query_threads |
min(CPUs, 4) | DuckDB threads for daemon aggregate queries; zero keeps the default |
query_temp_limit |
2GB |
Maximum spill-to-disk size for daemon aggregate queries; a query that spills past it fails with a DuckDB out-of-memory error |
If a Web UI query runs out of memory or temporary disk space, its error names
the query-limit settings above. Try filters that narrow the results. On the
machine running msgvault, check available memory and free disk space before
raising these limits in config.toml. Restart the daemon to apply the change,
then retry the query. The limits cap resource use; they do not reserve memory
or disk space. Cache builds have separate builder_* limits.
Start with the builder defaults, including at most two threads. More threads
can increase peak memory use. DuckDB's memory budget covers its buffer manager;
native allocations and Go memory add to the process total. A 24GB budget
therefore does not guarantee that a build stays below 24 GB of resident memory.
Leave room for the daemon, syncs, other applications, and those allocations.
For a constrained host, lower builder_memory_limit and set
builder_threads = 1 before raising the memory budget. See
DuckDB's memory guidance.
Builders spill temporary work under the cache staging directory, beside the
analytics cache. Check free space on that filesystem before increasing
builder_temp_limit. The default allows up to 32GB of spill in addition to
the existing cache and the new generation being staged. When DuckDB's SQLite
scanner is unavailable, the CSV fallback also writes a temporary copy of the
source tables beside the database, or in the system temporary directory if that
fails. A larger disk budget can help a large archive finish with a smaller
memory budget; it does not reserve space.
The relationship cache stores one fact per message and its direct sender or
recipient edges. Group-chat builds join the full member roster once per logical
conversation entry. Earlier messages retain their direct edges and owner
presence for relationship scores. Large groups therefore keep every member in
people searches without multiplying every message by the whole roster.
Direct chats and meetings retain their per-message participant attribution.
The [activity].max_direct_counterparts broadcast threshold continues to govern
the activity projection; changing it does not change relationship membership.
The daemon starts HTTP health and API routing before analytics cache
maintenance. With engine = "duckdb", analytics remain unavailable until a
usable cache is ready. If no usable cache can be built or opened, msgvault serve
fails instead of silently falling back. A failed automatic refresh after a sync
keeps serving the last usable publication, and the daemon checks again 15
minutes later even if later syncs write nothing. With auto_build_cache = false, use
msgvault build-cache, query --fresh, or sync --build-cache for explicit
cache maintenance. Deprecated in 0.17.0:
per-command analytics flags such as msgvault tui --force-sql,
msgvault mcp --force-sql, msgvault tui --no-cache-build, and
--no-sqlite-scanner were replaced by this daemon-level section. Use
engine = "sql" to force live SQL.
min_rebuild_interval limits automatic refreshes requested by syncs, queries,
and daemon startup. A restart or sync within the interval leaves the existing
publication in service and schedules a rebuild check when the interval ends.
A busy archive can therefore serve Parquet analytics that lag SQLite by
approximately the configured interval plus cache build time. Deleted messages
can remain visible in analytics query results until the next cache publication,
including while the interval has not elapsed and while a build runs. A zero
interval still allows stale rows during the build. Use query --fresh to wait
for analytics that include deletions committed before the request.
Explicit refresh requests and recovery of an absent, interrupted, incompatible, or otherwise unusable cache are not delayed.
The daemon runs automatic rebuilds in the background, outside the scheduled
sync that requested them, so syncs keep their cadence while a build runs. A
sync that finishes during a build does not discard it: the build publishes
one consistent read snapshot with that snapshot's counters and message boundary.
The next build appends later messages and repairs journaled child-row changes.
Participant links can refresh relationship data while retaining existing
message shards, including when new messages arrive in the same sync. Changes
to baked message facts, account identities, deletions, and failed syncs still
require a full rebuild. Archives without the child-row repair journal use a
conservative full rebuild after an overlapping sync. The published snapshot
remains usable across daemon restarts, and automatic follow-up builds still
honor min_rebuild_interval.
Cache build memory and temporary disk usage scale with archive size, so a
minimum interval can prevent repeated archive-scale work when sources sync
frequently. Changes under [analytics] take effect after the daemon restarts.
This setting governs the aggregate views (Senders/Domains/Labels/Time) and is ignored entirely when [data].database_url points at PostgreSQL — a PostgreSQL backend always uses live SQL for those views, and build-cache refuses to run against it. It does not affect the Web UI's Explore, Files, or People/domains workspaces, which require the SQLite + DuckDB/Parquet cache regardless of this setting and are unavailable on PostgreSQL; see PostgreSQL Backend for the current scope.
Default settings for msgvault backup. See Backup for the
capture, verify, and restore workflow.
| Key | Default | Description |
|---|---|---|
repo |
— | Default backup repository directory used when a backup subcommand omits --repo |
zstd_level |
0 |
Compression level for backup pack files. 0 uses msgvault's built-in default; otherwise use 1 through 19 |
When set, archive-access CLI commands use the remote server by default. Without [remote].url, they use the local background daemon instead. Pass --local to use the local daemon instead of the configured remote.
| Key | Default | Description |
|---|---|---|
url |
— | Remote API base URL (e.g. http://nas-ip:8080) |
api_key |
— | API key used by remote commands |
api_key_file |
— | Owner-only file holding the remote API key |
api_key_env |
— | Name of an environment variable holding the remote API key |
allow_insecure |
false |
Allow HTTP remote connections |
Affected CLI commands include search (FTS mode), query, show-message, stats, list-accounts, list-senders, list-domains, list-labels, identity subcommands, collection subcommands, export-eml, export-attachment, export-attachments, and tui. With a read-only remote client key, list-senders, list-domains, list-labels, query, vector search, tui, and mcp aren't available.
The same settings route mcp to a remote daemon. Secret precedence and file
requirements match server credentials. --local ignores the remote
destination and its secret sources. An unused destination's secret is not read.
Runtime keys from files or environment variables are never copied into saved
api_key fields by setup or export-token.
Scheduled sync sources for the web server. Each [[accounts]] entry defines a
cron schedule for automatic background syncing. Gmail, IMAP, Microsoft Teams,
and Discord sources are supported. For IMAP or Teams, use the account display
name/email when available rather than a raw provider identifier. Discord
schedules must use the exact guild ID because guild display names are mutable
and may be duplicated.
| Key | Default | Description |
|---|---|---|
email |
(required) | Account identifier/display name, or exact Discord guild ID, to sync |
schedule |
— | Cron expression for sync schedule (e.g., 0 * * * *) |
enabled |
true |
Whether scheduled sync is active for this account |
For example, schedule one previously registered Discord guild independently:
[[accounts]]
email = "123456789012345678"
schedule = "*/30 * * * *"
enabled = trueScheduled SMS Backup & Restore sources are configured with [[synctech_sms.sources]] entries. These are created automatically by msgvault add-synctech-sms-drive, but can also be edited directly.
| Key | Default | Description |
|---|---|---|
name |
(required) | Source name used by sync-synctech-sms <name> and scheduler logs |
enabled |
true |
Whether the source is active |
backend |
local |
local for a path on disk, or drive for Google Drive |
path |
— | Local XML/ZIP file or directory when backend = "local" |
folder_id |
— | Google Drive folder ID when backend = "drive" |
google_account |
— | Google account used for Drive access |
owner_phone |
(required) | Owner phone number in E.164 format |
schedule |
— | Cron expression used by msgvault serve |
include_sms |
true |
Import SMS records |
include_mms |
true |
Import MMS records |
include_calls |
true |
Import call logs |
include_attachments |
true |
Import MMS attachments |
stable_after |
10m |
How long Drive files must remain unchanged before import |
oauth_app |
— | Named Google OAuth app to use |
Scheduled Google Calendar sync is configured with top-level [[gcal]] entries. Each entry is one OAuth account; msgvault serve runs it on the given cron schedule (first run full-syncs and registers calendars, later runs are incremental). Authorize the account first with msgvault add-calendar.
[[gcal]]
name = "primary" # optional; defaults to email
email = "you@gmail.com" # OAuth account = token key
oauth_app = "" # optional named OAuth app
calendars = [] # optional calendarId filter; empty = owner+writer
schedule = "0 */6 * * *" # 5-field cron, no seconds
enabled = true
write_calendars = [] # explicit IDs; empty denies every event write
invite_calendars = [] # subset allowed to change guests or notify them
# calendar_aliases = { team = "team@example.com" }| Key | Default | Description |
|---|---|---|
name |
Source name used by sync-calendar <name> and scheduler logs |
|
email |
(required) | Google account that owns the token (the token key) |
oauth_app |
— | Named Google OAuth app to use |
calendars |
— | Specific calendar IDs to sync; empty syncs owned/writable calendars |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Enable the source for scheduled sync and live control; scheduling also requires schedule |
write_calendars |
[] |
Exact live calendar IDs allowed for event writes; empty denies all writes |
invite_calendars |
[] |
Exact IDs allowed to change guests, respond, or notify them; writes still require write_calendars |
calendar_aliases |
{} |
Names mapped to exact calendar IDs for live control; aliases do not expand permissions |
Event control also requires write consent.
calendars selects sync targets; it does not grant write authority. email selects
the OAuth token, while write_calendars selects the calendar that owns the event.
For example, person@example.com can create on team@example.com when Google
currently reports owner or writer for that calendar. primary resolves to the
live primary calendar ID before the daemon checks policy. List the actual ID in
both permission lists; neither list supports wildcards or alias names.
Archive joined rooms from native Matrix accounts. One
block controls every account registered with msgvault add-matrix; credentials
never belong in config.toml.
[matrix]
enabled = true # gate for the daemon schedule
schedule = "*/30 * * * *" # 5-field cron; empty = manual sync only
rooms = [] # exact room-ID include filter (empty = all)
exclude_rooms = [] # exact room IDs to skip; wins over rooms| Key | Default | Description |
|---|---|---|
enabled |
false |
Whether the daemon schedules Matrix sync |
schedule |
— | Cron expression used by msgvault serve |
rooms |
all joined rooms | Exact Matrix room IDs to sync |
exclude_rooms |
— | Exact Matrix room IDs to skip; exclusions win over inclusions |
Archive chats from a locally running Beeper Desktop. A single
block (not a list): the Beeper Desktop API is loopback-only, so there is one
instance per machine and the daemon must run beside it. Authorize first with
msgvault add-beeper.
[beeper]
# url = "http://localhost:23373" # Beeper Desktop API (default)
enabled = true # gate for the daemon schedule
schedule = "*/30 * * * *" # 5-field cron; empty = manual sync only
accounts = [] # accountID include filter (empty = all)
exclude_accounts = [] # skip networks archived natively, e.g. ["whatsapp"]
rate_limit_qps = 20 # request rate against the local API
media = true # download attachment bytes
media_scope = "all" # all, direct, or none
media_max_participants = 20 # skip media from larger rooms; 0 = no cap
max_media_mb = 250 # per-attachment download cap (MiB)
# [beeper.accounts_config.signal] # per-account override, keyed by accountID
# media = true
# max_media_mb = 500| Key | Default | Description |
|---|---|---|
url |
http://localhost:23373 |
Beeper Desktop API base URL |
enabled |
false |
Whether the daemon schedules Beeper sync |
schedule |
— | Cron expression used by msgvault serve |
accounts |
all | Beeper accountIDs to sync (include filter) |
exclude_accounts |
— | Beeper accountIDs to skip (wins over accounts) |
rate_limit_qps |
20 |
Request rate limit against the local API |
media |
true |
Download attachment bytes (failed downloads retry via backfill-beeper-media) |
media_scope |
all |
all, direct, or none; see Media policy |
media_max_participants |
20 |
Skip media from conversations above this many participants; 0 = no cap |
max_media_mb |
250 |
Per-attachment download cap in MiB (over-cap media is recorded as a size_cap skip and retried only after the cap changes) |
accounts_config |
— | Per-accountID media and max_media_mb overrides |
The daemon can send stored WAV and MP3 audio from any captured source, including messaging and email imports, to a separately running Docbank media service. The service needs Docbank's media HTTP routes. See Send audio to Docbank for the capture and processing rules.
[integrations.docbank]
enabled = true
url = "http://127.0.0.1:8080" # your Docbank daemon; the port is an example
api_key_env = "DOCBANK_API_KEY" # daemon environment variable with the key
all_sources_upload_consent = true # allow audio from every captured source to leave msgvault
# asr_profile = "asr" # optional Docbank profile for audio without source text| Key | Default | Description |
|---|---|---|
enabled |
false |
Schedule the stored-media job in msgvault serve |
url |
— | Docbank base URL: HTTPS, or HTTP on a loopback address. User info, query strings and fragments are rejected |
api_key_env |
— | Name of the daemon environment variable that holds the Docbank API key. It is read for each request and sent as X-Api-Key |
api_key |
— | Inline Docbank API key; takes priority over file and environment sources |
api_key_file |
— | Owner-only Docbank key file, read for each request; takes priority over api_key_env |
all_sources_upload_consent |
false |
Allow stored audio and explicit source transcripts from every captured source, including future providers, to be sent to url. Without it the job only records local state |
asr_profile |
— | Optional Docbank processing profile for stored audio without usable source text. An empty value retains audio without requesting processing. Msgvault rejects supplied-transcript, which Docbank reserves for supplied transcript input. |
The former Beeper-only upload_consent setting no longer enables uploads.
Existing users must explicitly set all_sources_upload_consent = true to resume
sending audio, including Beeper recordings.
The daemon reads these settings at startup, so restart it after a change. A new
url starts a separate delivery record; earlier rows stay. Disabling the route
stops the job and keeps its rows. A failed setup, such as an invalid url,
does the same and logs a warning. all_sources_upload_consent covers transport only; the
Docbank daemon's processing consent still decides whether a configured profile
may run. The route inspects stored CAS bytes, so MIME claims do not expand
Docbank's WAV and MP3 capability. Capture gaps and unsupported formats remain
typed local states.
Archive Slack workspaces. A single block covers every
registered workspace (tokens are per-workspace files). Authorize each
workspace first with msgvault add-slack.
[slack]
enabled = true # gate for the daemon schedule
schedule = "*/30 * * * *" # 5-field cron; empty = manual sync only
channels = [] # channel-name include filter (empty = all available)
exclude_channels = [] # channel names to skip, e.g. ["noise"]
private_channels = true # sync private channels
dms = true # sync one-to-one direct messages
group_dms = true # sync group direct messages
media = true # download shared-file bytes
media_scope = "all" # all, direct, or none
media_max_participants = 20 # skip files from larger channels; 0 = no cap
max_media_mb = 250 # per-file download cap (MiB)
# [slack.accounts_config.T0123456] # per-workspace override, keyed by team ID
# media = false| Key | Default | Description |
|---|---|---|
enabled |
false |
Whether the daemon schedules Slack sync |
schedule |
— | Cron expression used by msgvault serve |
channels |
all | Channel names to sync (include filter; never applies to DMs or group DMs) |
exclude_channels |
— | Channel names to skip (wins over channels) |
private_channels |
true |
Sync private channels; false pauses them without removing archived messages or affecting DMs |
dms |
true |
Sync one-to-one DMs; false pauses them without removing archived messages |
group_dms |
true |
Sync group DMs; false pauses them without removing archived messages |
media |
true |
Download shared-file bytes (failed downloads retry via backfill-slack-media) |
media_scope |
all |
all, direct (DMs and group DMs only), or none; see Media policy |
media_max_participants |
20 |
Skip files from conversations above this many members; 0 = no cap |
max_media_mb |
250 |
Per-file download cap in MiB (over-cap files are recorded as a size_cap skip and retried only after the cap changes) |
accounts_config |
— | Per-team-ID media and max_media_mb overrides |
For public channels only, set private_channels, dms, and group_dms to
false. These settings select what sync archives; the Slack token determines
what it can access. See Slack permissions
for a token restricted to public channels. A restricted token lists all public
channels, including unjoined ones; broader tokens list your memberships.
Media policy for Microsoft Teams chats and channels. Teams
sync itself is scheduled through [[accounts]]; this table only decides which
attachments are downloaded.
[teams]
media = true
media_scope = "all"
media_max_participants = 20
max_media_mb = 250
[teams.accounts_config."user@example.com"]
media = true
max_media_mb = 500| Key | Default | Description |
|---|---|---|
media |
true |
Download attachment and inline hosted-content bytes (failed downloads retry via backfill-teams-media) |
media_scope |
all |
all, direct (chats only, not channels), or none; see Media policy |
media_max_participants |
20 |
Skip media from chats and channels above this many members; 0 = no cap |
max_media_mb |
250 |
Per-attachment download cap in MiB |
accounts_config |
— | Per-account overrides of media and max_media_mb, keyed by the Teams account email |
Granola meeting-notes sync is configured with top-level [[granola]] entries.
Each entry is one Granola account. identifier is a stable source label;
account_email is the primary identity used for organizer attribution.
msgvault serve runs it on the given cron schedule. Register the account
first with msgvault add-granola. See
Meeting Transcripts.
[[granola]]
identifier = "work" # stable label; defaults to "default" for a single entry
account_email = "you@example.com" # required primary account identity
api_key = "grn_..." # from the desktop app's settings (Business plan)
schedule = "0 */6 * * *" # 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-granola <identifier> and scheduler logs |
account_email |
(required) | Normalized primary account identity used for is_from_me |
api_key |
(required) | Granola API key (grn_…) |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Config loading preserves identifier and rejects labels without an effective
email, instructing you to add account_email. msgvault add-granola confirms
the primary identity even if aliases already exist. Manage aliases with
msgvault identity add <identifier> <email>, then run
msgvault sync-granola <identifier> --full after identity changes to repair
existing meeting attribution. A scheduled source must still be registered in
the archive; removing it prevents the scheduler from silently recreating it.
Configure one top-level [[plaud]] entry per Plaud cloud account. Browser
OAuth stores credentials separately from this file. Enable Cloud Sync and
transcription in Plaud before syncing. See the
meeting guide for setup and preservation rules.
[[plaud]]
identifier = "work"
account_email = "you@example.com"
schedule = "30 */6 * * *"
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default for one unnamed entry |
Stable command, source, and token label; must be unique and contain no path separators, control characters, or surrounding whitespace |
account_email |
Required | Explicit account email; normalized to lowercase and checked against live Plaud identity |
endpoint |
https://mcp.plaud.ai/mcp |
MCP resource endpoint override; changing it requires new authorization |
schedule |
— | Five-field cron expression used by msgvault serve |
enabled |
false |
Whether a scheduled entry runs in the daemon |
Authorize with msgvault add-plaud <identifier> on the daemon host. The callback
uses localhost:8091/callback/plaud; tokens use
tokens/plaud_<identifier>.json. An existing source retains its confirmed
owner even if configuration changes. Use a new identifier for another account.
A scheduled entry must be registered; source removal prevents sync from
recreating it automatically.
Top-level [[twenty]] entries connect Twenty workspaces' call recordings through
read-only API keys. Keep these entries on the daemon host. Register each source
with msgvault add-twenty before syncing or scheduling it. See
Meeting Transcripts for API permissions.
[[twenty]]
identifier = "work"
account_email = "you@example.com"
base_url = "https://api.twenty.com"
api_key = "YOUR_READ_ONLY_API_KEY"
schedule = "15 */6 * * *"
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Stable label for commands and the twenty:<identifier> scheduler job |
account_email |
(required) | Actual primary email for account identity and organizer attribution |
base_url |
(required) | API root: https://api.twenty.com for Cloud, or the self-hosted instance origin |
api_key |
(required at registration/sync) | API key with read access to recordings, calendar events, and participants |
schedule |
— | Five-field cron expression used by msgvault serve |
enabled |
false |
Opt into daemon scheduling; manual sync remains available |
Origins require HTTPS except loopback HTTP and cannot include credentials,
queries, fragments, or path prefixes. Redirects are rejected. Identifiers must
be unique ignoring case; each entry requires account_email. Enabled sources
without a schedule are not scheduled. Removing a registered source prevents
scheduled sync from silently recreating it.
Circleback meeting sync is configured with top-level [[circleback]]
entries. Authentication is browser OAuth (msgvault add-circleback); no
secret lives in the config file. See
Meeting Transcripts.
[[circleback]]
identifier = "work" # stable label/token key; defaults to "default" for one entry
account_email = "you@example.com" # required primary account identity
schedule = "30 */6 * * *" # 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-circleback <identifier>, the token filename, and scheduler logs |
account_email |
(required) | Normalized primary account identity used for is_from_me |
endpoint |
production | MCP endpoint override (testing only) |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Config loading and alias repair follow the same rules as Granola: preserve the
label, add account_email, manage aliases with msgvault identity, and run
msgvault sync-circleback <identifier> --full after identity changes.
Circleback OAuth always confirms the primary identity; there is no identity
opt-out flag.
Notion meeting sync uses one top-level [[notion_meetings]] entry per Notion
identity. The meeting token needs AI Meeting Notes and Read Content access.
A personal access token (PAT) can read its user's meetings but cannot list users
or retrieve other users. To resolve Notion attendee IDs, configure a separate
internal integration with Read user information including email addresses.
The integration must belong to the same workspace.
[[notion_meetings]]
identifier = "notion-personal" # stable source label; defaults to "default" for one entry
account_email = "you@example.com" # required primary account identity
token = "ntn_..." # meeting token; keep this file private
users_token = "ntn_..." # optional workspace users integration token
schedule = "15 */6 * * *" # optional 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by commands and scheduler logs |
account_email |
(required) | Normalized primary identity for relationships; it is not assumed to be the meeting organizer |
token |
(required) | Meeting token; PAT or integration with Meeting Notes and Read Content access |
users_token |
— | Optional workspace integration token with Read user information including email addresses |
schedule |
— | Cron expression used by msgvault serve |
enabled |
false |
Whether the source is daemon-scheduled |
Manual and scheduled sync use the same optional users token from config. Token values stay out of diagnostics and archived evidence. See Notion attendee emails for lookup timeouts, failures, and retries.
Run msgvault add-notion-meetings <identifier> to validate access and register
the source before enabling a schedule. Removing the source prevents the
scheduler from recreating it. See Meeting Transcripts for
the 50-result discovery limit, attendee visibility, transcript retries, and
stored data.
Unreleased: configure one [[twilio]] entry per Twilio account or subaccount.
See the Twilio meeting guide for what sync stores.
[[twilio]]
identifier = "work"
account_email = "you@example.com"
account_sid = "AC00000000000000000000000000000001"
api_key_sid = "SK00000000000000000000000000000001"
api_key_secret = "your-key-secret"
region = "us1"
enabled = true
schedule = "15 */6 * * *"| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Stable source label; required with several entries |
account_email |
Required | Your primary identity; not treated as a caller |
account_sid |
Required | Account or subaccount SID |
api_key_sid, api_key_secret |
— | API key credentials, set together |
auth_token |
— | Account auth token instead of an API key |
region |
us1 |
us1, ie1 or au1; credentials must belong to that region |
intelligence_service_sid |
— | Only read Conversation Intelligence transcripts from this service; without it, transcripts from every service are read |
enabled |
false |
Allow daemon scheduling |
schedule |
— | Five-field cron expression |
media |
true |
Download recordings; false archives calls and transcripts only. After turning it back on, run sync-twilio --full to fetch the skipped recordings |
max_media_mb |
250 |
Per-recording size cap in MiB; 0 uses the default |
Muesli meeting sync uses one top-level [[muesli]] entry per Muesli database.
Configure these entries on the Mac where Muesli records. With [remote]
configured, the native client reads the local database and Contacts, then sends
selected meetings using the existing remote API credentials. Without remote
mode, the daemon reads the files on its own host. Remote Muesli sync and the
native completion hook are available on main and are not yet released.
[[muesli]]
identifier = "mac" # stable source label; defaults to "default" for one entry
account_email = "you@example.com" # required; you, the person who records
db_path = "~/Library/Application Support/Muesli/muesli.db" # optional; this is the default
phone_country_code = "1" # optional; convert national-format Contacts phones
schedule = "*/30 * * * *" # optional 5-field cron, no seconds
enabled = true| Key | Default | Description |
|---|---|---|
identifier |
default (single entry) |
Source name used by sync-muesli <identifier> and scheduler logs |
account_email |
(required) | Normalized primary identity; attributed as the organizer of every meeting |
db_path |
~/Library/Application Support/Muesli/muesli.db |
Muesli database path; ~ expands, and a relative path resolves against the config directory when --config is used |
contacts |
true |
Resolve attendees through Apple Contacts; the process reading the files needs Full Disk Access |
contacts_path |
~/Library/Application Support/AddressBook |
Apple Contacts data folder; expands like db_path |
phone_country_code |
— | Country calling code (1–3 digits, such as "1" or "44") used for Contacts phone numbers typed without one; unset means only international numbers are used |
schedule |
— | 5-field cron used by the local daemon or recorder-side sync-muesli --watch |
enabled |
false |
Whether automatic rescanning includes this source |
Run msgvault add-muesli <identifier> to check the database and register the
source before enabling a schedule. On a remote recorder, keep
msgvault sync-muesli --watch running to use enabled schedules. Run
msgvault sync-muesli <identifier> --full
after identity changes to repair existing meeting attribution. Removing the
source prevents the scheduler from recreating it. See
Meeting Transcripts for what gets stored.
Top-level toggle and backend marker for semantic/hybrid search. SQLite vector search requires a build with sqlite_vec support (default via make build). PostgreSQL vector search requires a build with the pgvector tag and a PostgreSQL [data].database_url. See Vector Search for prerequisites, initial embedding, and the full workflow.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on vector and hybrid search. When false, mode=vector and mode=hybrid return vector_not_enabled. |
backend |
sqlite-vec |
Backend marker. Supported values are sqlite-vec and pgvector; the concrete backend is selected from [data].database_url. |
db_path |
<data_dir>/vectors.db |
SQLite vector database path. Ignored by the PostgreSQL pgvector backend. |
skip_extension_create |
false |
PostgreSQL only. Skip CREATE EXTENSION IF NOT EXISTS vector when pgvector is already installed by an administrator. |
External OpenAI-compatible embedding endpoint used to convert message text into vectors. msgvault does not host a model; it calls the endpoint you configure. Use a local or self-hosted endpoint (Ollama, llama.cpp server, LM Studio, etc.) when message text must stay on your machine or network. Hosted endpoints also work but receive the text being embedded.
| Key | Default | Description |
|---|---|---|
api_format |
openai |
Request contract: openai (OpenAI-compatible /embeddings, one vector per message chunk) or voyage-contextual (Voyage /contextualizedembeddings; pins model = "voyage-context-4" and embeds chat conversation windows and turn-aware meeting chunks as contextual documents). |
endpoint |
(required) | HTTP(S) base URL for an OpenAI-compatible embeddings API. msgvault appends /embeddings (for example, set http://localhost:11434/v1, not .../embeddings). |
model |
(required) | Model name or serving alias to pass in each request (e.g., nomic-embed-text). Use a new alias when the served weights or vector recipe change; msgvault does not verify the checkpoint behind an alias. |
dimension |
(required) | Expected vector width. Responses with a different width are rejected. The OpenAI-compatible client does not send a dimensions field or slice returned vectors; configure the server to return the desired width. |
document_prefix |
"" |
Model-specific instruction prepended to every document chunk after chunking (for example, "search_document: " for nomic-embed-text). The prefix does not reduce max_input_chars; maximum 4096 UTF-8 bytes. |
query_prefix |
"" |
Model-specific instruction prepended to every vector-search query (for example, "search_query: " for nomic-embed-text); maximum 4096 UTF-8 bytes. |
api_key_env |
— | Name of an environment variable containing the API key. Omit for anonymous endpoints. |
batch_size |
32 |
Embedding inputs per HTTP call. Long messages can contribute multiple chunk inputs. |
timeout |
30s |
Per-request timeout. |
max_retries |
3 |
Retries per batch on transient failures. |
max_input_chars |
32768 |
Character cap per embedding chunk, counted in characters rather than tokens. Too high and chunks are rejected or silently truncated; too low and long messages split into more chunks, adding embedding overhead. For example, start around 6000 for a 2k-token model such as Ollama's nomic-embed-text, then check representative content. See Matching max_input_chars to your embedder's context window. |
eta_window |
10 |
Number of recent progress samples used for ETA smoothing. |
For a native 768-dimensional EmbeddingGemma 2 setup, see the EmbeddingGemma 2 text example.
Instead of naming an environment variable in api_key_env, you can store a
provider API key through Settings in the Web UI or the TUI. Stored keys live in
tokens/provider-credentials.json under the data directory with owner-only
file permissions. They are never written to config.toml. After saving,
Settings shows only a masked hint of the key, its first three and last three
characters, and whether it comes from the store or from the environment; the
key itself is never returned.
A stored key takes precedence over the environment variable named by
api_key_env. Each stored key is bound to the endpoint origin (scheme, host,
and port) it was saved for. If you later change the endpoint to another origin,
the stored key is removed automatically and must be entered again, so a key
is never sent to a host it was not entered for.
Changing a stored key for vector or multimodal (visual) embeddings requires a daemon
restart, like the other [vector] settings. Person enrichment and sweep keys
apply on the next run.
On unreleased main, a stored person-enrichment suppression key also takes
precedence over a custom suppression_key_env. Previously, a custom variable
won. Before upgrading an installation that has both, ensure the stored key
matches the variable's value. If existing suppression records were made with
a different key, enrichment stops with ErrSuppressionKeyMismatch.
On unreleased main, the host owner can also install keys without the Web UI:
msgvault credentials set vector.embeddings --from-file /run/secrets/embedding-key
msgvault credentials set vector.multimodal --stdin < /run/secrets/visual-key
msgvault credentials set people.enrichment/research --endpoint https://api.example.com/search --stdin
msgvault credentials set people.enrichment/suppression --stdin < /run/secrets/suppression-key
msgvault credentials list --json
msgvault credentials import-envThe CLI lists IDs and bound origins without values. File input follows the
server secret-file rules. Standard input is trimmed and limited to
64 KiB. Suppression keys must contain at least 32 bytes after trimming;
set and import-env reject shorter values before saving them.
--endpoint defaults to the configured provider endpoint; suppression
keys have no endpoint. import-env copies present configured vector,
multimodal, named enrichment, and suppression variables once, preserving
existing stored keys. These host commands do not start the daemon or change
provider consent. Run them on the daemon host; --local selects a local home
when remote access is configured.
Sweep providers use their own named profile credentials. Use
msgvault person provider add --api-key-stdin or --credential-env for them;
the legacy people.sweep store ID is not consumed by current sweep runs.
The index generation fingerprint includes the model, dimension, document and query prefixes, preprocessing settings, max_input_chars, embedding policy, and scope. Changing those settings triggers a stale-index error on the next vector/hybrid query. For an existing account-scoped generation built with CLI flags, set matching [vector.embed.scope].accounts and restart the daemon; otherwise run msgvault embeddings build --full-rebuild.
Controls text normalization before embedding.
| Key | Default | Description |
|---|---|---|
strip_quotes |
true |
Drop quoted reply blocks (> ... lines, reply preambles) before embedding. |
strip_signatures |
true |
Drop trailing signature blocks (content after -- ). |
strip_html |
true |
Convert HTML-only bodies to text and remove HTML markup before embedding. |
strip_base64 |
true |
Remove base64/data blobs before HTML stripping so encoded data does not crowd out prose. |
strip_url_tracking |
true |
Remove common tracking parameters such as utm_*, fbclid, and gclid from URLs. |
collapse_whitespace |
true |
Normalize repeated horizontal whitespace and blank lines. |
Hybrid ranking parameters applied at query time.
| Key | Default | Description |
|---|---|---|
rrf_k |
60 |
Reciprocal Rank Fusion constant. Higher values flatten score differences between signals. |
k_per_signal |
100 |
Candidate pool size drawn from each signal (BM25 or vector) before fusion. |
subject_boost |
2.0 |
Multiplier applied when a query term matches a message's subject line. |
max_page_size_hybrid |
50 |
Hard cap on page_size for vector/hybrid responses. Set to 0 to disable clamping. |
sqlite_accelerator |
auto |
Use a ready SQLite approximate index. Set to exact to keep exhaustive vector search. PostgreSQL ignores this setting. |
ann_nprobe |
8 |
SQLite index partitions searched per query. Higher values trade latency for recall. |
ann_oversample |
8 |
Approximate candidates requested per result before exact reranking. Range: 1–128. |
ann_threads |
CPU count, max 128 |
Native worker threads used by msgvault embeddings optimize. Range: 1–128. |
Accelerator tuning does not change the embedding generation fingerprint. It changes how stored vectors are searched or optimized, not how text is sent to the embedding provider.
Optional scope for newly built embedding generations. The zero value embeds the
full archive. A scoped generation embeds only matching messages.message_type
values:
[vector.embed.scope]
message_types = ["teams"]Scoped generations are intentionally partial. Vector and hybrid queries against
a scoped index must include a compatible message_type filter, such as
msgvault search "release planning" --mode hybrid --message-type teams; an
unscoped vector/hybrid query returns index_scope_mismatch instead of using the
partial index as if it covered the full archive.
| Key | Default | Description |
|---|---|---|
message_types |
[] (all types) |
Embed only messages of these types. |
accounts |
[] (all accounts) |
Embed only these accounts' messages, by canonical account identifier (display names are rejected here — they are not stable identities for a privacy boundary). Resolved to source IDs at startup; an unknown identifier fails vector initialization (or the CLI command). The daemon's scheduled embeds honor this scope, so it also acts as a privacy boundary: unlisted accounts' text is never sent to the embedding endpoint. |
accounts and message_types compose (both filters apply). The CLI flags
--account/--collection on msgvault embeddings build/resume override
accounts for a single run. Either scope dimension is part of the generation
fingerprint: changing it requires msgvault embeddings build --full-rebuild,
and because the fingerprint records archive-local source IDs, re-adding an
account under a new source ID also requires a rebuild. Account-scoped indexes
do not gate search the way message-type scopes do: out-of-scope accounts
simply have no vector matches and rank on BM25 alone in hybrid mode.
Optional background scheduling for the embed worker inside msgvault serve.
Empty config disables scheduled embedding; you can still run
msgvault embeddings build by hand.
| Key | Default | Description |
|---|---|---|
cron |
— | 5-field cron expression. Empty string disables the standalone cron. |
run_after_sync |
false |
Run an embed pass after successful scheduled Gmail, IMAP, Teams, and Discord syncs. Other sources use the standalone cron or a manual build. |
msgvault setup providers supplies run_after_sync = true and
cron = "*/15 * * * *" when it enables a text lane. It preserves either key
when explicitly set, including false and "". Already enabled text lanes keep
their schedules. See
Recommended Configuration.
Semantic people search: one curated document per durable person, built from
searchable non-sensitive attributes, embedded into the text-search generation.
Requires [vector] enabled = true and a separate consent
(msgvault person provider consent --semantic-embeddings --yes).
| Key | Default | Description |
|---|---|---|
enabled |
false |
Embed curated person documents and serve msgvault person search. |
retention_posture |
— | Your assertion about the embedding provider's retention; must be explicit (not unknown). |
training_posture |
— | Your assertion about the embedding provider's training use; must be explicit. |
Independently consented visual attachment lane over Voyage. Every value has a
default except the probe manifest; uploads fail closed without it, and a
daemon started with enabled = true and no manifest refuses every vector
lane until one exists.
| Key | Default | Description |
|---|---|---|
enabled |
false |
Turn on the visual lane. |
provider |
voyage |
Only legal value. |
endpoint |
https://api.voyageai.com/v1 |
Pinned provider root; other origins are refused. |
api_key_env |
VOYAGE_API_KEY |
Environment variable holding the key. A key alone enables nothing. |
model |
voyage-multimodal-3.5 |
Pinned model. |
dimension |
1024 |
Pinned dimension. |
capabilities_file |
— | Manifest written by msgvault multimodal probe --seeds <dir> --out <file> --yes. |
max_context_chars |
4000 |
Owning-message text sent with each attachment. |
include_images |
true |
Embed still images (JPEG, PNG, WebP). |
include_animated_gifs |
false |
Embed animated GIFs; requires include_images and a manifest that authorized them. |
include_video |
true |
Embed direct-input MP4 video. |
allow_image_queries |
true |
Allow multimodal search --image. |
[vector.multimodal.scope] accepts the same message_types and accounts
keys as [vector.embed.scope]; [vector.multimodal.schedule] accepts the
same cron and run_after_sync keys as [vector.embed.schedule]. Consent is
recorded per generation by msgvault multimodal build --yes.
Dated activity projection and per-person contact state (first and last
contact, inbound/outbound, interaction count, inferred channel). It is the
deterministic source of "when did we last talk" for every person and runs
hourly by default inside msgvault serve. msgvault activity build runs it
by hand; --backstop rescans the whole archive.
Scheduled projection commits at most ten batches per pass. When other scheduled work has waited for a minute, it stops after its current batch. A pass always stops at two minutes. A pass with committed progress releases the operation gate and resumes behind queued work without waiting for the next cron tick. Reaching the two-minute limit before any batch commits records an error and waits for the next scheduled or manual trigger, avoiding repeated retries of the same batch.
Identity reconciliation and timezone or max_direct_counterparts changes save
their progress in the archive. A new identity revision restarts identity
reconciliation; completed batches remain committed. Manual builds are not
limited to ten batches or two minutes.
| Key | Default | Description |
|---|---|---|
schedule |
17 * * * * |
5-field cron used by msgvault serve. Empty disables the scheduled job. |
timezone |
UTC |
IANA zone name for day bucketing. Local is rejected because the projection keys replay on the persisted zone name. |
max_direct_counterparts |
25 |
Largest conversation still projected as direct activity between its participants. |
batch_size |
500 |
Messages per projection batch. |
By default, msgvault stores everything under ~/.msgvault (macOS/Linux) or C:\Users\<you>\.msgvault (Windows). To use a different location, you have two options:
--home flag (per-command):
msgvault sync --home /mnt/data/msgvaultMSGVAULT_HOME environment variable (persistent):
export MSGVAULT_HOME=/mnt/data/msgvaultBoth options are equivalent: config.toml is loaded from the specified directory, and all data (database, tokens, attachments) is stored there. The --home flag takes priority over MSGVAULT_HOME.
The home or [data].data_dir directory may be a symlink to an existing
directory. Local daemon bookkeeping resolves the symlink and applies its
ownership and permission checks to the target directory.
| Variable | Description |
|---|---|
MSGVAULT_HOME |
Base directory for all data (default: ~/.msgvault) |
MSGVAULT_BIND_ADDR |
Server bind address or iface:NAME; serve --bind wins |
MSGVAULT_API_PORT |
Server port from 0 to 65535; serve --port wins |
MSGVAULT_API_KEY |
Inline server key for this process |
MSGVAULT_API_KEY_FILE |
Mounted server key file |
MSGVAULT_API_KEY_ENV |
Name of the environment variable holding the server key |
MSGVAULT_ALLOW_INSECURE |
Allow unauthenticated non-loopback serving |
MSGVAULT_BACKUP_REPO |
Backup repository path |
MSGVAULT_CORS_ORIGINS |
Comma-separated browser origins; empty clears the list |
MSGVAULT_CORS_CREDENTIALS |
Whether browser CORS requests may use credentials |
MSGVAULT_TRUSTED_PROXIES |
Comma-separated proxy IP addresses or CIDRs; empty clears the list |
MSGVAULT_REMOTE_URL |
Remote daemon URL for all commands with remote support |
MSGVAULT_REMOTE_API_KEY |
Inline remote key for this process; export-token --api-key wins |
MSGVAULT_REMOTE_API_KEY_FILE |
Mounted remote key file |
MSGVAULT_REMOTE_API_KEY_ENV |
Name of the environment variable holding the remote key |
MSGVAULT_REMOTE_ALLOW_INSECURE |
Allow plaintext HTTP to the remote daemon |
MSGVAULT_TELEMETRY_ENABLED |
Set to 0 to turn off anonymous telemetry; any value overrides [telemetry] enabled (Telemetry) |
These runtime controls are available on unreleased main. Environment values
override TOML. For each server or remote credential group, setting any of its
three environment variables selects that group's environment sources; within
the group, inline wins over file, then named environment. A supplied empty
credential variable is an error when that destination is used. Booleans accept
Go's strconv.ParseBool values, such as true, false, 1, and 0; empty or
invalid values fail. Origin and proxy lists trim whitespace and ignore empty
entries, including a trailing comma. Explicit export-token --to, --api-key,
and --allow-insecure choices are saved even when they match an environment
override; environment-only values are not saved. An explicit
--allow-insecure=false overrides the environment and saved configuration,
requires HTTPS, and saves false after a successful export. The setup wizard also saves
new choices that match an override; keeping existing settings leaves them
unchanged on disk.
For example, a supervised daemon can start without a config file:
MSGVAULT_HOME=/data MSGVAULT_BIND_ADDR=0.0.0.0 MSGVAULT_API_PORT=8080 msgvault servePersist /data across restarts so the archive and generated key survive. The
published stock image is ghcr.io/kenn-io/msgvault; it runs as UID/GID 1000
and stores its home at /data. Use an image containing these unreleased
features once published. No startup hook or entrypoint wrapper is required.
msgvault serve sends anonymous usage telemetry to PostHog:
daemon_activewhen the daemon starts, then once on each later UTC day while it runs.app_openedwhen the web UI opens, then on its first window focus on a later UTC day. The browser reports it to the daemon, never to PostHog.screen_viewedwhen a web or terminal screen opens, counted once per installation per UTC day across both interfaces and daemon restarts. The daemon keeps the current day's screens intelemetry-screen-views.jsonbeside its install ID.
For app_opened, the web UI records the day it last reported in browser storage, which the
browser keeps separately for each daemon address. With the default
api_port = 0, the daemon picks a new port each time it starts, so the web UI
reports again after a daemon restart. Tabs that open together, or a browser
that blocks storage, can also each send one. Each event carries only:
- the product name and source (
msgvault,daemon) - on
app_opened, the surface (web) - on
screen_viewed, the first surface (webortui) and a fixed screen name:everything,directory,directory_review,files,operations,relationships,saved_views,sources,deletions,settings,message,email,texts, ormeetings. Unknown names are dropped. - the msgvault version and commit
- the operating system and CPU architecture
- a random install ID kept in
telemetry-install.jsonin the data directory, and the whole hours since it was created - metadata the PostHog Go library adds itself: library name and version (
$lib,$lib_version), OS name, Go version, and where available the OS version and distribution
Events exclude message content, contact records, account or source identifiers, filenames, and search text. They ask PostHog to skip person profiles and location lookup. The daemon queues each event and sends it in the background, so an event can be lost if the network is down or the daemon stops first.
When telemetry is on, msgvault serve says so in its startup output and log.
To turn it off, set this in config.toml and restart a running daemon:
[telemetry]
enabled = falseMSGVAULT_TELEMETRY_ENABLED in the environment that starts the daemon
overrides the config: 0, false, no or off turns telemetry off, and any
other value turns it on. TELEMETRY_ENABLED=0 also turns it off. A CLI command
that starts a local daemon passes its environment to it. Builds made with the
kit_posthog_disabled tag never send telemetry.
The default home is ~/.msgvault on macOS/Linux and C:\Users\<you>\.msgvault
on Windows. It is created automatically; existing directory permissions are
left unchanged. Configuration stays under the home unless --config selects
another file. The data paths below use [data].data_dir, which defaults to the
home; [log].dir can override the log location.
| File | Description |
|---|---|
<home>/config.toml |
Configuration file |
msgvault.db |
SQLite database (system of record when PostgreSQL is not configured) |
attachments/ |
Content-addressed attachment files |
tokens/ |
OAuth tokens and stored provider credentials |
tokens/server-api-key |
Persisted daemon API key, reused on later loopback and non-loopback starts |
logs/ |
Structured log files (when file logging is enabled) |
analytics/ |
Parquet cache files for Web UI and TUI analytical views |
telemetry-install.json |
Random anonymous install ID for telemetry; created only while telemetry is on |
telemetry-screen-views.json |
Current UTC day's screen claims shared by web and terminal UIs; created only while telemetry is on |
Copy only the sections you need and replace example paths and credentials.
[data]
# Base data directory (default: ~/.msgvault)
data_dir = "/path/to/msgvault/data"
# User-requested exports (default: {data_dir}/exports)
export_dir = "/path/to/msgvault/exports"
# Database URL (default: {data_dir}/msgvault.db; PostgreSQL DSN supported)
database_url = "/path/to/msgvault.db"
# Keep attachment content as individual files instead of creating packs.
# loose_attachments = true
[oauth]
# Path to Google OAuth client secrets JSON for browser OAuth
client_secrets = "/path/to/client_secret.json"
# Google service account key for Workspace domain-wide delegation (optional)
# service_account_key = "/path/to/service-account.json"
# Named OAuth apps for Google Workspace orgs (optional)
[oauth.apps.acme]
client_secrets = "/path/to/acme_workspace_secret.json"
# service_account_key = "/path/to/acme_service_account.json"
[microsoft]
# Azure AD app registration client ID (required for M365)
client_id = "your-azure-app-client-id"
# redirect_uri = "http://localhost:8089/callback/microsoft" # default
# tenant_id = "your-tenant-id" # optional, default "common"
# Optional source-scoped Fastmail alias inventory.
[[fastmail]]
source_id = 14
api_token = "replace-with-a-Fastmail-API-token"
auto_confirm_identities = false
[discord]
# Per-attachment download cap (default: 50 MiB)
max_media_bytes = 52428800
# Skip attachments from rooms with more than this many participants
# (default: 20; 0 = no cap). Shared by [beeper], [slack], and [teams].
media_max_participants = 20
# Trailing edit/delete/reaction repair window (default: seven days)
edit_rescan_window = "168h"
[discord.guilds."123456789012345678"]
# Channel, thread, and forum-post IDs; empty include means all accessible.
include = ["456789012345678901"]
exclude = ["567890123456789012"]
[log]
# Persistent structured file logging (opt-in)
enabled = true
# dir = "/path/to/logs" # default: <data_dir>/logs
# level = "info" # debug, info, warn, error
# sql_trace = false # log every SQL query (verbose)
# sql_slow_ms = 100 # slow query threshold in ms
[sync]
# Gmail API rate limit (requests per second)
rate_limit_qps = 5
[server]
# API server settings (used by `msgvault serve` and `msgvault daemon`)
# api_port is optional; omit it (or set 0) to auto-select an open port that
# clients discover automatically. Set a fixed port for remote/NAS deployments.
api_port = 0
bind_addr = "127.0.0.1"
api_key = "your-secret-key"
daemon_idle_timeout = "20m" # background daemon idle timeout; "0s" disables
daemon_auto_restart = "newer" # newer, never, or always
daemon_auto_start = true # false when a supervisor runs msgvault serve
[analytics]
# Daemon-side analytics engine for Web UI, TUI, and aggregate HTTP views:
# "auto" starts on live SQL and switches to DuckDB after cache maintenance.
# "sql" always uses live SQL. "duckdb" requires a usable Parquet cache.
engine = "auto"
# Build a stale/missing cache during daemon startup and after scheduled syncs.
auto_build_cache = true
# Minimum age of a usable cache before a scheduled sync may rebuild it again.
# min_rebuild_interval = "6h"
[backup]
# Default repository for `msgvault backup`.
repo = "~/Backups/msgvault"
zstd_level = 0
[deletion]
# Durable consent for remote deletion execution. Opt in deliberately;
# defaults to false.
remote_enabled = false
[remote]
# Remote msgvault endpoint for CLI remote mode
url = "http://nas-ip:8080"
api_key = "remote-api-key"
allow_insecure = true
# Scheduled sync accounts
[[accounts]]
email = "you@gmail.com"
schedule = "0 * * * *"
enabled = true
[vector]
# Semantic and hybrid search (opt-in)
enabled = true
backend = "sqlite-vec"
# backend = "pgvector" # with a PostgreSQL database_url and pgvector build
[vector.embeddings]
endpoint = "http://localhost:11434/v1"
model = "nomic-embed-text"
dimension = 768
document_prefix = "search_document: "
query_prefix = "search_query: "
eta_window = 10
[vector.preprocess]
strip_quotes = true
strip_signatures = true
strip_html = true
strip_base64 = true
strip_url_tracking = true
collapse_whitespace = true
[vector.embed.scope]
# Empty means embed the full archive. Set this for partial generations.
message_types = ["sms", "mms"]
# Use stable account identifiers, not numeric source IDs. This keeps a scoped
# generation usable after a daemon restart.
# accounts = ["you@work.example"]
[attachments.documents]
# Hosted extraction is opt-in and requires a separately recorded consent.
enabled = false
provider = "mistral"
region = "eu"
api_key_env = "MISTRAL_API_KEY"
model = "mistral-ocr-4-0"
retention_posture = "zdr"
training_posture = "opted-out"
max_file_bytes = 52428800
max_pages_per_document = 500
max_response_bytes = 67108864
max_normalized_chars = 25000000
max_spool_bytes = 536870912
min_free_space_bytes = 1073741824
request_timeout = "5m"
max_retries = 3
max_pages_per_run = 10000
max_estimated_cost_usd_per_run = 50
# Set both pricing fields together to include a cost estimate in manual build preflight.
# estimated_cost_usd_per_1000_units = 0.001
# pricing_assumption_on = "2026-08-17"
[attachments.documents.scope]
# Empty includes every supported message type.
message_types = ["email"]
[attachments.documents.index]
lexical = true
store_chunk_text = true
[[synctech_sms.sources]]
name = "phone-backups"
enabled = true
backend = "drive"
folder_id = "google-drive-folder-id"
google_account = "you@gmail.com"
owner_phone = "+14155551234"
schedule = "30 4 * * *"
[telemetry]
enabled = true # false turns off anonymous usage telemetry