A federated search system for Cooklang recipes that allows decentralized publishing and centralized discovery through RSS/Atom feeds.
- 🔍 Unified search with powerful query syntax powered by Tantivy
- 📡 RSS/Atom feed crawler with automatic updates
- 🏷️ Advanced filtering by tags, ingredients, time, difficulty, and more
- 🌐 Web UI for browsing and searching recipes
- 💻 CLI tools for searching, downloading, and publishing recipes
- 🔄 Background scheduler for automated feed crawling
- 🛡️ Rate limiting to protect API endpoints from abuse
- 🐳 Docker support for easy deployment
# Clone the repository
git clone <repository-url>
cd federation
# Start the server
docker-compose up -d
# Access the web UI
open http://localhost:3000- Rust 1.75 or later
- SQLite 3
# Clone the repository
git clone <repository-url>
cd federation
# Copy environment variables
cp .env.example .env
# Download Tailwind CSS CLI
./scripts/download-tailwind.sh
# Start development server (with Tailwind watch mode)
./scripts/dev.shThe server will be available at http://localhost:3001
# Build Tailwind CSS once
./tailwindcss -i ./styles/input.css -o ./src/web/static/css/output.css
# Run database migrations
cargo run -- migrate
# Start the server
cargo run -- serveThe federation search supports powerful query syntax powered by Tantivy's QueryParser:
# Search all fields
breakfast
# Search specific field
tags:breakfast
title:pasta
ingredients:tomato
difficulty:easy# Boolean operators
pasta AND tags:italian
breakfast OR brunch
# Exclusion
chocolate -tags:dessert
# Range queries
total_time:[0 TO 30] # 30 minutes or less
servings:[4 TO 8] # Serves 4-8 people
# Complex combinations
chocolate tags:dessert difficulty:easy
pasta AND tags:italian AND total_time:[0 TO 30]Use quotes for multi-word field values:
tags:"quick breakfast"
title:"chocolate chip cookies"title- Recipe titlesummary- Recipe descriptioninstructions- Cooking instructionsingredients- Ingredient listtags- Recipe tagsdifficulty- Difficulty level (easy, medium, hard)servings- Number of servingstotal_time- Total cooking time in minutesfile_path- Source file path (for GitHub recipes)
# Basic search
cargo run -- search "chocolate cookies"
# Field-specific search
cargo run -- search "tags:breakfast"
# Complex query
cargo run -- search "pasta AND tags:italian AND total_time:[0 TO 30]"cargo run -- download 123 --output ./recipes# Generate an Atom feed from .cook files
cargo run -- publish --input ./my-recipes --output feed.xmlDetect and store the locale of recipes that don't have one yet (author-declared
locale: metadata wins, otherwise the language is detected from the recipe text).
Touched recipes are re-indexed automatically.
# Tag only recipes that don't have a locale yet
cargo run -- backfill-locales
# Recompute the locale for every recipe, including ones already tagged
cargo run -- backfill-locales --forceFix titles of GitHub recipes indexed by earlier versions (declared title:
metadata, else a readable version of the file name) and delete recipes whose
content duplicates an older one (the same file at two paths, a fork, a mirror
feed). The search index is updated in step. Safe to rerun.
cargo run -- cleanupThe API is read-only JSON. Errors are {"error": "message"} with a 4xx/5xx
status, including for query strings that do not parse (page=abc, or a
parameter given twice: list filters are comma-separated). /api/* is
rate-limited per client IP (see API_RATE_LIMIT); a limited request gets
429 with a Retry-After header (seconds).
GET /health- Health checkGET /ready- Readiness checkGET /api/stats- Totals:total_recipes,total_feeds,total_tags,total_ingredients,active_feeds
-
GET /api/search- Search recipes. Every term inqmust match (vegan tags:dessertis vegan and dessert); useORfor alternatives and-to exclude. Words are stemmed, socakefinds "cakes".Param Example Meaning qq=pasta tags:italianQuery string (field syntax as above) localelocale=deLanguage; enalso matchesen-UStagstags=vegan,dessertHas all tags (stemmed, case-insensitive) include_ingredientsinclude_ingredients=garlic,lemonUses all ingredients exclude_ingredientsexclude_ingredients=peanutUses none of them max_timemax_time=30Total time ≤ minutes min_servings,max_servingsmin_servings=2&max_servings=6Inclusive range difficultydifficulty=easyExact, case-insensitive feed_idfeed_id=12Only this feed sortsort=newestrelevance(default) ornewestpage,limitpage=2&limit=20Paging (limit max 100) Invalid numbers, an unknown
sortor a malformedqreturn400. Each result card hasid,title,summary,tags,locale,total_time_minutes,servings,difficulty,image_urlandfeed: {id, title}. Any field exceptid,titleandtagsmay be null. -
GET /api/facets?tag_limit=200- Tag, language and difficulty values with recipe counts, for filter UIs. Cached for 5 minutes;tag_limitdefaults to 200, max 1000.
GET /api/recipes/:id- Recipe details, includinglocale(e.g."de","en-US") andlocale_source("declared"if set via a Cooklanglocale:key, or"detected"if inferred from the recipe text)GET /api/recipes/:id/download- Download .cook file
GET /api/feeds- List feeds (page,limit,status)GET /api/feeds/:id- Get feed details
Feeds are registered through config/feeds.yaml, not the API.
Environment variables (see .env.example):
| Variable | Description | Default |
|---|---|---|
DATABASE_URL |
Database connection string | sqlite:./data/federation.db |
HOST |
Server host | 0.0.0.0 |
PORT |
Server port | 3000 |
EXTERNAL_URL |
External URL for CLI | http://localhost:3000 |
API_RATE_LIMIT |
API requests per second | 100 |
CRAWLER_INTERVAL |
Seconds between feed updates | 3600 |
MAX_FEED_SIZE |
Maximum feed size in bytes | 5242880 (5MB) |
MAX_RECIPE_SIZE |
Maximum recipe size in bytes | 1048576 (1MB) |
RATE_LIMIT |
Crawler requests per second per domain | 1 |
INDEX_PATH |
Search index directory | ./data/index |
RUST_LOG |
Logging level | info,federation=debug |
This release adds feed_id, indexed_at, image_url and feed_title to the
Tantivy schema and makes servings and total_time indexed, for the new
structured search filters, sort=newest and richer result cards. As with
earlier schema changes, the server refuses to start against an index built
by a previous version.
Rebuild before starting serve:
rm -rf data/index # or your configured INDEX_PATH
federation backfill-locales --forceWith Docker Compose, stop the app first, then run the rebuild in a one-off container:
docker compose stop app
rm -rf data/index
docker compose run --rm app federation backfill-locales --force
docker compose up -d appbackfill-locales --force is the full rebuild: it re-indexes every recipe with
its tags, ingredients and feed title, fills missing servings, total time
and difficulty from the recipe's Cooklang metadata, and fills the ingredient
list of every recipe that has none stored from its Cooklang content (recipes
that already have ingredients keep them). (federation reindex <url> is a
different command: it deletes one feed's recipes from the database and
re-crawls that feed, and it does not rebuild the search index.) Recipes with
no stored content are not indexed by the rebuild and will not appear in
search results or facet counts until their content is fetched again.
Ship this whole branch as one release: production only needs one index rebuild, not one per commit.
Other behaviour changes:
- Feed recipes are indexed as they are crawled. Before, recipes from
RSS/Atom feeds (and their
<category>tags) reached the search index only throughbackfill-locales. Existing feed recipes reach the index through the rebuild above; nothing further is needed for them. - Servings, total time and difficulty come from Cooklang metadata.
GitHub and feed recipes now store
servings,time(orprep time+cook time) anddifficultyfrom their.cookmetadata, so themax_time,min_servings/max_servingsanddifficultyfilters and the difficulty facet also match them. Before, those columns stayed empty for GitHub recipes, and for feed recipes unless the feed entry set them. The values live in the database, not only in the index, so existing recipes get them from the full rebuild above:backfill-locales --forcefills empty columns from each recipe's stored content, writes them to the database and indexes them. No re-crawl is needed. A re-crawl would not help anyway, because the GitHub indexer skips files whose SHA has not changed. - Feed recipes store their ingredients. Before, only GitHub recipes had
an ingredient list, so the
include_ingredients/exclude_ingredientsfilters silently ignored feed recipes (exclude_ingredients=peanutstill returned feed recipes with peanuts). The crawler now stores ingredients from each.cookfile the same way the GitHub indexer does, and keeps the stored list when an updated file fails to parse. Existing feed recipes get their ingredient lists from the full rebuild above:backfill-locales --forcefills them from each recipe's stored content, writes them to the database and indexes them. No re-crawl is needed. - Rate limiting is per client and matches
API_RATE_LIMIT. It used to be one bucket for everyone that refilled one request everyAPI_RATE_LIMITseconds. Now each client getsAPI_RATE_LIMITrequests per second with bursts of twice that. If you run behind a reverse proxy, the proxy must connect to this server from a loopback or private address, and it must overwrite (not append to) theX-Forwarded-Forheader with the real client address — e.g. in nginx,proxy_set_header X-Forwarded-For $remote_addr;. Otherwise the first hop ofX-Forwarded-Foris trusted as the client address, and a client can set that header itself to pick its own rate-limit bucket. - A malformed
qreturns400with the parser's message instead of500. - The website search form has the same filters as the API.
The text analyzer changed (stemming, plain-text instructions), which changes the Tantivy schema. Delete the index, rebuild it, then repair recipes indexed by the previous GitHub indexer:
rm -rf data/index # or your configured INDEX_PATH
federation backfill-locales --force
federation cleanup--force re-indexes every recipe, not only those without a locale. cleanup
rewrites slug-style titles from the recipes' own metadata and removes
duplicate recipes, and keeps the index in step.
This release adds a locale field to the Tantivy search schema (used to tag
and filter recipes by language). Tantivy pins field ids to the schema stored
on disk, so a search index built before this change is incompatible — the
server now refuses to start against a mismatched index, with an error
telling you what to do.
Before deploying this version, delete the existing index and rebuild it:
rm -rf data/index # or your configured INDEX_PATH
federation backfill-locales --forcebackfill-locales runs pending database migrations and re-indexes every
recipe it touches; with --force that is every recipe, so this single command
rebuilds the search index and backfills locales in one step. Run it before
starting serve again.
To build the project for production:
# Build everything (CSS + Rust binary)
./scripts/build.shThis will:
- Build Tailwind CSS with minification
- Build the Rust binary in release mode
The output will be:
- Binary:
./target/release/federation - Minified CSS:
./src/web/static/css/output.css
To run in production:
# Set environment variables
export DATABASE_URL="sqlite:./data/federation.db"
export PORT=3000
# Run the server
./target/release/federation servecargo testcargo clippy -- -D warningsMigrations are located in the migrations/ directory and are automatically applied on server startup.
- Web Framework: Axum with Tokio async runtime
- Database: SQLite (PostgreSQL compatible)
- Search Engine: Tantivy full-text search
- Feed Parsing: feed-rs for RSS/Atom
- Recipe Parsing: cooklang-rs
- Templates: Askama with Tailwind CSS
- CLI: Clap for command-line interface
- Rate Limiting: tower-governor for API protection
For production use, it's recommended to deploy this service behind a reverse proxy (nginx, Caddy, Traefik, etc.) that:
- Terminates TLS/SSL
- Sets proper
X-Forwarded-Forheaders for accurate rate limiting - Provides additional DDoS protection
- Handles load balancing if running multiple instances
Rate limiting works best when proper IP information is available via reverse proxy headers.
To make your recipes discoverable:
- Create
.cookfiles in a directory - Generate an Atom feed:
cargo run -- publish --input ./recipes --output feed.xml
- Host the feed and .cook files at a public URL
- Add your feed to a federation server:
curl -X POST http://localhost:3000/api/feeds \ -H "Content-Type: application/json" \ -d '{"url": "https://your-site.com/feed.xml"}'
See LICENSE file for details.
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.