Skip to content

Add age-based retention for cached artifacts - #428

Open
pinguinfuss wants to merge 3 commits into
git-pkgs:mainfrom
pinguinfuss:issue-306-artifact-retention
Open

pinguinfuss wants to merge 3 commits into
git-pkgs:mainfrom
pinguinfuss:issue-306-artifact-retention

Conversation

@pinguinfuss

Copy link
Copy Markdown
Contributor

Infrastructure for #306. This adds a storage.retention block (default, ecosystems, packages, sweep_interval) and a sweep that evicts artifacts nobody has downloaded for longer than the configured time.

No ecosystem is enabled yet, so this PR changes nothing on its own. Each ecosystem will get its own small PR that registers it from its handler. Until then, naming an ecosystem or package in the config fails validation, and default doesn't apply to it. On startup the proxy logs which ecosystems retention covers.

What's in here:

  • an artifact expires when both fetched_at and last_accessed_at are older than the cutoff
  • the sweep scans by id window, flushes buffered hits first and clears at most 10,000 artifacts per run; files go through pending_deletes
  • sweep_interval must be at least a minute
  • reclaim keeps working through due entries for up to 30s per tick and stops after a failed delete
  • new index on pending_deletes(queued_at) (migration 010)
  • proxy_artifacts_evicted_total{reason, ecosystem} for LRU and retention
  • gradle.build_cache.max_age accepts "7d"

Tested on SQLite and Postgres 16. The new database tests run against both, including one with a non-UTC local zone.

Cached artifacts stayed forever unless storage.max_size pushed them out.
storage.retention adds a default, per-ecosystem and per-package duration
after which an artifact nobody downloaded gets evicted.

Ecosystems opt in one at a time: a handler registers with
retention.Default, and until it does, configuring that ecosystem (or one
of its packages) fails validation, and the default doesn't touch it. This
commit registers none, so nothing changes yet. Startup logs which
ecosystems the default covers and which it doesn't.

An artifact counts as expired when both its fetch time and its last
access are older than the cutoff. Checking only the last access would
evict a refetched artifact again right away, since clearing a record
keeps its old access time and a cache miss doesn't record a hit.

The sweep scans artifacts in id windows, so no single query holds the
SQLite connection for long, and checks each row against its own rule.
Expired records are cleared and their files go through pending_deletes
instead of being deleted directly, because a buffered hit can make an
artifact that's being served right now look expired. Buffered hits are
flushed before each sweep, and one sweep clears at most 10,000
artifacts.

Since a big first sweep can queue a lot, reclaim now keeps working
through due entries for up to 30 seconds per tick instead of stopping
after 100, and pending_deletes gets an index on queued_at.

Also:
- proxy_artifacts_evicted_total{reason, ecosystem} counts LRU and
  retention evictions.
- gradle.build_cache.max_age accepts "7d", as its comment always said.

Refs git-pkgs#306
Describes the new storage.retention block, the order rules apply in,
that the default only covers ecosystems that support retention yet, and
what to expect from the first sweep over an old cache. Also lists the
new eviction counter and notes that gradle max_age takes days.

Refs git-pkgs#306
On Postgres, fetched_at and last_accessed_at are TIMESTAMP columns
without a zone. They hold the local time the proxy wrote, but lib/pq
reads them back labelled UTC. East of UTC that makes every artifact look
hours younger than it is, so the sweep evicted late: two hours in
Berlin, nine in Tokyo. The candidates' times are now read back as local
time, and the database tests run on both SQLite and Postgres, with one
that sets the local zone to Asia/Tokyo.

A few smaller things:

- When a delete fails, reclaim stops for the rest of the minute instead
  of trying every queued path. Otherwise a storage outage turns into
  thirty seconds of failed deletes and warnings every minute.
- If the buffered hits can't be written, the sweep skips that round.
  Running on stale access times could clear something that is being
  downloaded right now.
- DefaultCanonical only accepts a package key when rebuilding it gives
  the same PURL. For alpine and deb it would otherwise produce keys
  that never match.
- sweep_interval has to be at least a minute.
- New tests cover the eviction counter and an artifact that gets a hit
  between the scan and the clear.
- The docs no longer say nothing changes without retention, and they
  point out that the example's ecosystem and package rules only pass
  validation once those ecosystems support retention.

Refs git-pkgs#306
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant