# ytj-23signs

`ytj.23signs.org`: a friendlier search interface over Finland's YTJ/PRH Trade Register, built to
serve **both** human users (a search box) and AI agents (a JSON API, OpenAPI spec and the usual
discovery documents for tool-calling).

Live at `https://ytj.23signs.org` (Cloudflare Worker, custom domain in `wrangler.toml`). The data
is the full Trade Register, held in Cloudflare D1 and refreshed daily from PRH's open data dump.

## Data

PRH publishes the Trade Register two ways under the YTJ open data v3 API
(`https://avoindata.prh.fi/opendata-ytj-api/v3`, keyless, CC BY 4.0, credit PRH):

- a paginated `/companies` endpoint (100 records per page), which reports about 827,000
  entities because it includes historical and removed registrations;
- a daily `all_companies` zip (about 97 MB compressed, 1.4 GB of JSON) with the current
  register: 464,751 companies at the time of writing, 462,131 of them registered and 2,620
  unregistered or pending. It is refreshed around 04:45 UTC.

The service is built on the dump. One download a day replaces hundreds of thousands of paginated
calls, and the paginated API is only used for spot checks and for picking up registrations made
since the last dump.

**Scope limitation (inherited from PRH):** the register data does not cover `toiminimi` (sole
proprietorships) that are not registered in the Trade Register. `?businessId=3645785-4`, for
example, returns nothing because that ID belongs to a toiminimi. Any agent or human using this
should expect a "not found" for valid-looking IDs that are unregistered sole proprietorships.

## Architecture

```
PRH all_companies.zip  --fetch-->  crawler/state/ytj.sqlite  --push-->  Cloudflare D1  <--  src/worker.js (Hono)
     (daily, ~04:45 UTC)           (local source of truth,              (companies, names,       |
                                    one content hash per company)        names_fts trigram)      +-- public/ (UI, docs)
```

- **`src/worker.js`**: Hono app. Three API routes (`/api/v1/companies`,
  `/api/v1/companies/:businessId`, `/api/v1/dataset`), the discovery documents under
  `/.well-known/`, and content negotiation on `/`. Every JSON body carries `_links`, every error
  is RFC 9457 problem+json.
- **`src/data.js`**: one data interface, two backends. When the `DB` binding exists (production)
  it queries D1: exact `business_id` lookups, and name search through an FTS5 trigram index over
  current and auxiliary names, 25 results per page, bounded rows read (no total count). Without
  the binding (plain `wrangler dev`, tests, a fresh clone) it falls back to `src/data.json`, a
  bundled sample of 212 companies. Both return identically shaped company objects; the
  `dataset` object in responses tells you which one answered.
- **`crawler/ytj_crawl.py`**: fetches the dump, stream-parses it into a local SQLite, and pushes
  changed rows to D1 over the HTTP API within a daily write budget. See `crawler/README.md`.
- **`crawler/schema.sql`**: D1 schema (`companies`, `names`, `names_fts`, `meta`).

Company `status` stays two-valued (`active` | `not_currently_registered`) for existing clients;
`registerStatus` adds PRH's REK_KDI code list (`unregistered`, `registered`,
`removed_from_register`, `startup_not_registered`, `ceased`, `unknown`).

### Load status

The register goes into D1 in daily budgeted batches (see Cost below). Until the load finishes,
`GET /api/v1/dataset` reports `complete: false` and a `pending` count, and the same `dataset`
object is embedded in every search response and every 404, so a client can tell "not loaded
yet" from "not in the register". `SKILL.md` and `llms.txt` tell agents to check it.

## Tech stack

Hono on Cloudflare Workers with a D1 database, matching the convention used by `camara`, `fin`,
`alpha` and other apps under `~/code/apps/`: `src/worker.js` for the API, static assets in
`public/` for the frontend, `wrangler.toml` for config. `run_worker_first` routes `/`, `/api/*`,
`/.well-known/*` and the agent docs through the worker so they carry discovery headers;
everything else is served straight from `public/`.

## Structure

```
src/worker.js        Hono API: /api/v1/companies (search), /api/v1/companies/:businessId (lookup), /api/v1/dataset
src/data.js          Data layer: D1 backend (production) and bundled-sample fallback (local dev)
src/data.json        Bundled sample (212 PRH company records, raw v3 shape), local dev only
crawler/ytj_crawl.py PRH dump -> local SQLite -> D1 loader (uv run --script)
crawler/schema.sql   D1 schema
crawler/state/       Downloaded zip, local SQLite, logs (not committed)
public/index.html    Search UI
public/app.js        Frontend fetch + rendering logic
public/style.css     Styling
public/openapi.yaml  Machine-readable API spec (also openapi.json)
public/SKILL.md      Agent skill file; public/llms.txt, robots.txt, sitemap.xml, webmcp.js
```

## Agent discovery

Everything an agent needs to find and call the API without reading this file. Modelled on
`~/code/apps/camara`, minus auth (this API has none). Base: `https://ytj.23signs.org`.

| URL | What it is |
| --- | --- |
| `/openapi.yaml`, `/openapi.json` | OpenAPI 3.1 for the three endpoints, with `_links`, `dataset` and problem+json schemas |
| `/llms.txt` | llmstxt.org index: summary plus links to every document below |
| `/SKILL.md` | Agent skill: when to use, parameters, response and error shapes, load status, limitations, examples |
| `/.well-known/api-catalog` | RFC 9727 API catalog as an RFC 9264 linkset (`application/linkset+json`); `Accept: application/json` gives a plain JSON view |
| `/.well-known/ai-catalog.json` | ARD (Agentic Resource Discovery) manifest: host plus entries for the API, skill, agent card and tool card, with representative queries |
| `/.well-known/ard.json` | Same ARD manifest at the v0.9 spec path |
| `/.well-known/mcp/server-card.json` | MCP-style server card listing the two tools with JSON Schema inputs; honest `endpoint: null`, there is no hosted MCP transport |
| `/.well-known/agent-card.json` | A2A agent card: REST interface, no authentication, two read-only skills |
| `/.well-known/agent-skills/index.json` | Agent Skills index pointing at `/SKILL.md` |
| `/robots.txt` | Allow all, Content Signals (`ai-train=yes, search=yes, ai-input=yes`), Sitemap and Agentmap pointers |
| `/sitemap.xml` | Every URL in this table plus the dataset endpoint and two example API calls |
| `/webmcp.js` | WebMCP shim: registers `ytj_search_companies` and `ytj_get_company` for a browser agent on the landing page |
| `/README.md` | This file |

Every `/api/*` response and `/` carry a `Link` header (`rel="service-desc"` to the catalog
and `/openapi.yaml`, `rel="describedby"` to `/SKILL.md` and `/llms.txt`). Every JSON body
carries `_links` (`self`, `collection`, `search`, `openapi`, `catalog`, `describedby`, plus
`next` and `prev` on paginated searches); a company's `self` is its canonical
`/api/v1/companies/{businessId}`. Errors are RFC 9457 `application/problem+json` with `type`,
`title`, `status`, `detail`, `instance`; a 404 under `/api/` lists `availableEndpoints` and a
404 on a single lookup embeds `dataset`. `GET /` with `Accept: application/json` redirects to
the catalog; with `Accept: text/markdown` it returns `llms.txt`.

## Run locally

```bash
cd ~/code/apps/ytj-23signs
npm install
npm run dev        # wrangler dev, serves on http://localhost:8787
```

Plain `wrangler dev` has no D1 binding, so the API answers from the bundled 212-company sample
and `GET /api/v1/dataset` reports `backend: "sample"`. To exercise the real register locally,
run against the deployed D1 database:

```bash
npx wrangler dev --remote
```

Then open `http://localhost:8787` in a browser, or query the API directly:

```bash
curl "http://localhost:8787/api/v1/dataset"
curl "http://localhost:8787/api/v1/companies?businessId=0112038-9"
curl "http://localhost:8787/api/v1/companies?name=Nokia"
```

An AI agent can read `http://localhost:8787/openapi.yaml` (or `.json`) to learn the endpoints
and call them directly. Tests: `npm test` (vitest, sample backend).

Loading or refreshing the production database is the crawler's job; see `crawler/README.md`.

## Cost

Everything runs on Cloudflare's free tiers except the D1 write budget, which decides how fast
the register loads:

- **D1 Free**: 100,000 rows written per day. A company costs about 6 written rows (its row, one
  row per name, the index and the FTS trigram postings), so the crawler's default budget of
  95,000 rows pushes roughly 16,000 companies a day and the initial load of 464,751 companies
  takes about four weeks. After that, only companies whose content hash changed are pushed, and
  the daily delta fits comfortably.
- **Workers Paid** (USD 5/month): 50 million rows written per month included. Run the crawler
  with `--write-budget 0` and the whole register loads in one sitting (about 2.8 million rows).

Reads are cheap on both plans: a business ID lookup is one row, a name search is bounded at 26
companies plus the trigram index walk, and `dataset` is a tiny `meta` table read.
