# Coaster World DATA

Coaster World DATA is the independent Laravel service that centralizes canonical park, attraction, coaster, POI, calendar, live-status, queue-time, provenance and media data for Coaster World.

## Local environment

The current development environment is Laravel 13 / PHP 8.3+ / MySQL 8, with application timestamps stored in UTC. Park schedules use each park's IANA timezone.

The local Laragon URL is normally:

```text
http://coaster-world-data.test
```

Secrets and provider credentials must stay in `.env`; they must never be committed to the repository.

## Installation / update

After replacing the project files with a delivered archive, restore the local `.env`, then run:

```bat
php artisan optimize:clear
php artisan migrate:status
php artisan storage:link
```

`storage:link` only needs to be run if `php artisan about` reports `public\storage` as `NOT LINKED`.

## Reference data

The project includes idempotent seeders for the stable reference data currently required by DATA:

- configured source providers;
- park types;
- attraction types;
- POI types.

Run:

```bat
php artisan db:seed
```

The reference seeders can be run again safely. They do not create demo users and do not reactivate sources or taxonomy entries that an administrator has disabled. Existing source reliability overrides are also preserved.

To seed only the DATA references:

```bat
php artisan db:seed --class=Database\\Seeders\\ReferenceDataSeeder
```

## Data principles

- Internal relational IDs remain BIGINT values.
- Public entities use permanent ULIDs in `public_id`.
- External-provider IDs are stored in `external_mappings`, never as canonical primary IDs.
- Unknown values stay `NULL`; DATA must not invent zero values or fake dates.
- Raw provider observations remain distinct from canonical current status.
- Provider failure does not mean that a park or attraction is closed.
- Canonical data keeps evidence and conflict history through the provenance tables.
- User/community content remains outside DATA unless explicitly migrated later.

## Main schema blocks

- Reference data: organizations, resorts, parks, areas, manufacturers, ride models, attractions, coaster details and POIs.
- Provenance: sources, external mappings, fetches, field evidence and conflicts.
- Calendar: park days and local-time intervals.
- Live: raw observations, wait-time history, canonical park/attraction status and provider health.
- Media: media assets and polymorphic entity attachments.

## Tests

When the local PHP installation includes the extensions required by PHPUnit and SQLite, run:

```bat
php artisan test
```

The reference-data feature tests verify idempotency, ULID creation and preservation of administrator overrides.

## DATA construction mode

Before the canonical catalogue is populated, the server runs in a safety-first construction mode. The default environment values are:

```env
DATA_POPULATION_ENABLED=false
DATA_PROVIDER_SCHEDULES_ENABLED=false
DATA_CANONICAL_AUTO_APPLY_ENABLED=false
DATA_AI_AGENTS_ENABLED=false
```

With these defaults, provider dry-runs and operational health checks are allowed, but providers cannot create or change canonical park/attraction data, scheduled imports do not run, canonical proposals cannot be auto-applied, and AI agents cannot start. This prevents accidental population while the DATA platform is still being completed.

The `/admin/platform` screen exposes the current locks, the declared provider capability matrix, ingestion runs, canonical decisions and the AI-agent registry.

`provider_capabilities` describes what each configured source is intended to provide. `ingestion_runs` groups provider work, `data_proposals` stages source/agent suggestions without mutating canonical entities, and `canonical_decisions` records explainable decisions. AI agents are registered in `ai_agents` and remain disabled by default; later model drivers must create proposals rather than bypassing evidence and canonical review.

After applying the v0.4 platform migrations, run:

```bat
php artisan migrate
php artisan db:seed
php artisan optimize:clear
```

Do not enable `DATA_POPULATION_ENABLED` yet while the remaining DATA server blocks are still under construction.

Check the current safety locks and platform counters with:

```bat
php artisan cw:data:status
```

## ThemeParks.wiki catalogue sync

DATA includes its first automated provider integration for the ThemeParks.wiki V1 API. The catalogue sync imports destination/resort records and their parks, stores stable external IDs in `external_mappings`, records the raw provider response in `source_fetches`, keeps name evidence in `field_evidence`, and updates `provider_health`.

After migrations and reference seeding, test the provider without changing canonical entities:

```bat
php artisan cw:themeparks:sync-destinations --dry-run --limit=5
```

During DATA construction mode, stop after the dry-run. Real catalogue population is intentionally blocked while `DATA_POPULATION_ENABLED=false`.

Once park mappings are verified, DATA can import each park's ThemeParks.wiki child catalogue. This phase promotes only `ATTRACTION` entities into the canonical attraction table; shows and restaurants remain provider data until dedicated canonical models are introduced. Existing exact-name matches are created as pending mappings instead of being trusted automatically.

Test a few mapped parks without changing attractions or mappings:

```bat
php artisan cw:themeparks:sync-entities --dry-run --limit=5
```

During DATA construction mode, keep this command in `--dry-run` mode. Real entity population remains intentionally blocked until the platform foundation is complete.

The destination catalogue schedule (03:20 UTC) and attraction catalogue schedule (03:40 UTC) are registered but remain inactive until both `DATA_POPULATION_ENABLED=true` and `DATA_PROVIDER_SCHEDULES_ENABLED=true`. On a long-running local DATA instance, Laravel's scheduler can be kept active with:

```bat
php artisan schedule:work
```

In production, use Laravel's normal scheduler/cron setup instead of keeping an interactive terminal open.

`THEMEPARKS_WIKI_API_KEY` is optional for public endpoints. If a key is configured, it must remain in `.env`. ThemeParks.wiki usage and attribution requirements remain the responsibility of the deployment; DATA does not expose provider credentials or copy them into fetch logs.

## Administration graphique

À partir de `v0.3.0-dev.1`, Coaster World DATA dispose d'une première interface d'administration protégée.

Après application du patch :

```bash
php artisan optimize:clear
php artisan migrate
php artisan cw:admin:create
```

Puis ouvrir :

```text
http://coaster-world-data.test/admin
```

L'interface permet actuellement de consulter les indicateurs DATA, les resorts, les parcs, les attractions canoniques, la santé des fournisseurs, les collectes récentes, les mappings à vérifier et les conflits ouverts. Les fiches parc/attraction exposent leur provenance et leur état Live lorsqu'il existe. Les sources peuvent être activées/désactivées et la synchronisation ThemeParks.wiki des attractions peut être déclenchée depuis une fiche parc disposant d'un mapping vérifié.

Aucune inscription publique n'est proposée : un compte d'administration se crée uniquement en ligne de commande avec `cw:admin:create`.

## Multi-provider connectivity layer

Before any mass population of parks, attractions, coasters, POIs or Live data, DATA can now register and test the provider families planned for the platform without writing canonical entities.

Registered adapters:

- ThemeParks.wiki;
- Queue-Times;
- Wikidata;
- OpenStreetMap / Overpass;
- Wikipedia / Wikimedia;
- DATAtourisme;
- Foursquare Places;
- IGN / Géoplateforme;
- RCDB (manual reference only until explicit written permission allows automated reuse);
- official park/operator sources (template, configured individually later);
- internal manual administration.

Provider checks are intentionally independent from canonical population. A connectivity probe can update `provider_health` and append an auditable `source_fetches` row, but it cannot create or edit a park, attraction or POI.

After applying this patch and reseeding the stable reference definitions:

```bat
php artisan db:seed
php artisan optimize:clear
php artisan cw:data:status
```

Inspect provider configuration without network traffic:

```bat
php artisan cw:providers:check --all --no-network
```

Run the real connectivity probes explicitly:

```bat
php artisan cw:providers:check --all
```

Or test one provider:

```bat
php artisan cw:providers:check queue_times
```

The same probes are available from `/admin/platform`. Missing credentials are reported as configuration state rather than as provider downtime.

Credentials remain exclusively in `.env`. In particular, DATAtourisme requires `DATATOURISME_API_KEY`, while the current Foursquare Places integration expects `FOURSQUARE_SERVICE_KEY`. RCDB automated access is intentionally not implemented because its published terms require prior written permission to use its content to construct another database/application.

Keep all DATA construction locks disabled (`false`) during this infrastructure phase. Provider probes do not require `DATA_POPULATION_ENABLED=true`.

## Autonomous research missions

Coaster World DATA is designed so an administrator can ultimately enter a park or attraction and let DATA orchestrate the research workflow from a single mission.

The current construction phase introduces `/admin/research` and keeps real execution locked by default with:

```env
DATA_RESEARCH_EXECUTION_ENABLED=false
DATA_AI_AGENTS_ENABLED=false
DATA_CANONICAL_AUTO_APPLY_ENABLED=false
```

A mission can already be planned safely. DATA enumerates the provider capabilities and AI agents that would participate, records why each step is ready, blocked, manual or not yet implemented, and always ends with a separate canonical review step.

Restricted sources are never bypassed. In particular, RCDB capabilities marked as requiring permission remain blocked for autonomous use. Future AI agents may research the same facts from independent permitted sources, official pages or open data, but must not use automation to circumvent a source's usage restrictions.

The intended long-term workflow is:

```text
Target park / attraction
        ↓
Research mission
        ↓
Permitted providers + official sources + AI research agents
        ↓
Evidence / source fetches / proposals
        ↓
Conflict and quality analysis
        ↓
Canonical decision
        ↓
Coaster World DATA
```

During server construction, creating a mission does not populate the canonical park/attraction/POI catalogue.

## Enrichment orchestrator

`v0.4.0-dev.5` adds the completeness and cost strategy layer that sits between provider discovery and future AI execution.

DATA now maintains explicit enrichment profiles for parks, attractions and coasters. Each profile classifies expected encyclopedia fields as `required`, `recommended` or `optional`, declares which canonical paths satisfy the field, and records whether an AI fallback is allowed.

Provider-to-field coverage is stored separately in `provider_field_coverages`. A provider is only considered executable for a field when its capability is declared, connected, configured and allowed by the existing provider policy. Restricted providers remain visible in the plan but are never promoted to an automatic step.

The strategy is deliberately provider-first:

```text
Canonical entity / research target
        ↓
Completeness assessment
        ↓
Lowest-cost permitted provider coverage
        ↓
Re-assess missing fields
        ↓
AI fallback only for fields still missing
        ↓
Evidence / proposals / canonical review
```

Live status and queue-time data are intentionally excluded from this encyclopedia enrichment layer. They must continue to come from timestamped Live observations and must never be guessed by an AI agent.

After applying the patch:

```bat
php artisan migrate
php artisan db:seed
php artisan optimize:clear
php artisan cw:data:status
```

The administration page `/admin/enrichment` shows the three current completeness profiles and their provider coverage. Creating or replanning a research mission also generates a persisted completeness assessment and a field-level provider plan. AI steps are now true fallbacks: when a permitted provider can cover a missing field, the matching AI agent is marked as fallback rather than as the first source of data.

All DATA population, research execution, AI execution and canonical auto-apply locks remain unchanged and disabled by default during this construction phase.


## Targeted AI research runtime

`v0.4.0-dev.6` adds the execution layer used only after the enrichment orchestrator has exhausted lower-cost provider coverage for the requested fields.

The runtime remains safe by default:

```env
DATA_RESEARCH_EXECUTION_ENABLED=false
DATA_AI_AGENTS_ENABLED=false
DATA_AI_DRIVER=simulation
```

With these defaults, an administrator can simulate an AI step from `/admin/research` without making an external request, spending money, changing the research-step state or creating a factual proposal. Simulation exercises the runtime, audit log and budget snapshot only.

Each real AI run receives immutable execution limits for input/output tokens, web-search calls, cost, timeout and attempts. Runs execute on the dedicated `ai-research` Laravel queue, record an event timeline, usage, estimated cost and evidence, and can only stage `data_proposals`. They never write canonical park, attraction or coaster fields.

The OpenAI driver uses the Responses API with structured output and optional built-in web search. Web evidence is accepted only when its URL is returned by the search tool itself. Restricted domains configured by `DATA_AI_RESTRICTED_DOMAINS` are blocked in the search tool and rejected again before evidence is persisted. Web content is treated as untrusted evidence and cannot override runtime instructions.

During the DATA construction phase, keep the real execution gates locked and validate only simulations. Example CLI simulation for an existing AI research step ID:

```bat
php artisan cw:research:run-ai 123 --simulate
```

When the server foundation is later approved for controlled real tests, the deployment will require an API key in `.env`, a non-simulation driver, an enabled AI agent and a queue worker. Do not enable these yet while the canonical engine and remaining DATA platform milestones are unfinished.

Relevant optional runtime settings are documented in `.env.example`. Cost rates are configuration values rather than hard-coded business rules so they can be updated when provider pricing changes.

## Normalized provider collection

`v0.4.0-dev.9` adds the provider-normalization contract that sits between external APIs and the canonical pipeline. Provider responses can now be previewed in a common DATA format without writing anything to MySQL, then staged later behind an independent safety lock.

The default remains deliberately locked:

```env
DATA_PROVIDER_COLLECTION_ENABLED=false
```

A safe preview performs the external request and normalization in memory only. It does not create a `source_fetch`, normalized staging row, proposal or canonical value:

```bat
php artisan cw:providers:collect queue_times park_catalogue --limit=3
php artisan cw:providers:collect wikipedia encyclopedia --entity-type=park --query="Parc Astérix" --limit=3
php artisan cw:providers:collect wikidata entity_enrichment --entity-type=coaster --external-id=Q123
```

Connected collection adapters currently cover ThemeParks.wiki, Queue-Times, Wikidata, OpenStreetMap/Overpass, Wikipedia/Wikimedia, DATAtourisme, Foursquare Places and IGN/Géoplateforme. RCDB remains a restricted/manual reference source, while official-site and manual adapters stay special-purpose until concrete source definitions are registered.

When the staging lock is deliberately opened in a later validation phase, `--stage` persists the raw fetch plus provider-neutral records in `provider_normalized_records`. `--propose` additionally projects only already-mapped encyclopedia fields into `field_evidence` and pending `data_proposals`. Live-only fields remain in normalized staging and continue to use the dedicated Live pipeline rather than encyclopedia proposals.

Relation fields such as an attraction's park or manufacturer are never converted from a provider label directly into an internal `*_id`. A proposal is created only when the referenced external entity already has a verified `external_mapping`; otherwise the normalized relation remains staged for later mapping. This prevents textual provider IDs or names from being written into canonical foreign keys.

The normalizer layer is intentionally extensible. Adding a future provider normally requires:

1. a `ProviderAdapter` implementation;
2. `CollectsProviderData` when the provider can return normalized records;
3. registry registration;
4. a `sources` seed definition plus provider capabilities;
5. field-coverage definitions for fields the normalizer can actually emit;
6. credentials/endpoints in `.env` / `config/services.php` when required;
7. normalization tests.

No core schema change is normally needed for another provider. A migration is only expected when the new source introduces a genuinely new canonical entity, field family or platform capability that the current generic contracts cannot represent.

## Runtime H24

La couche runtime est conçue pour survivre aux redémarrages et limiter les effets d'un fournisseur instable sans écrire silencieusement dans le référentiel canonique.

Après les migrations de la version runtime :

```bash
php artisan cw:runtime:status
```

En développement Laragon/Windows, `tools/run_data_runtime_windows.bat` démarre deux fenêtres : le scheduler Laravel et le worker de queues `maintenance,provider,ai-research,default`. Elles doivent rester actives pendant les tests locaux. En production, utilisez un gestionnaire de processus/service adapté à l'hébergement plutôt qu'un terminal interactif.

Le runtime apporte : verrou anti-doublon sur les collectes provider, limitation locale par capacité, backoff progressif, circuit breaker après échecs répétés, reprise/fermeture des runs figés, heartbeat scheduler/worker, rétention des payloads RAW et conservation des jobs échoués pour diagnostic.

Les tâches de heartbeat et maintenance peuvent fonctionner pendant le mode construction : elles ne créent ni parc, ni attraction, ni valeur canonique. Les collectes provider restent contrôlées par `DATA_PROVIDER_COLLECTION_ENABLED` et les autres verrous DATA.

## DATA API v1 and security

`v0.4.0-dev.13` introduces the versioned read API that the Coaster World site and mobile application can consume later, without coupling them to DATA's internal MySQL identifiers.

Public canonical endpoints are available under `/api/v1`. They expose ULID `public_id` values only and never expose source credentials, internal provider payloads, canonical decision internals or numeric database IDs. Initial resources include resorts, parks, attractions/coasters, POIs, manufacturers, taxonomies, park calendars and the current canonical Live status.

Examples:

```text
GET /api/v1/parks
GET /api/v1/parks/{public_id}
GET /api/v1/parks/{public_id}/attractions
GET /api/v1/parks/{public_id}/calendar
GET /api/v1/attractions/{public_id}
GET /api/v1/live/attractions?park={park_public_id}
```

The catalogue API supports validated filters and bounded pagination. HTTP caching uses `Cache-Control` plus `ETag`; Live endpoints use a shorter TTL than encyclopedia/catalogue endpoints. API responses also include `X-Request-Id` and security headers. Errors use a stable envelope containing an error code, safe message and request ID.

The internal operations endpoint is deliberately separate:

```text
GET /api/v1/internal/status
```

It requires a bearer token (or `X-Data-Token`) configured only in `.env`. Generate a strong value locally with:

```bat
php artisan cw:api:token
```

Copy the generated value into:

```env
DATA_INTERNAL_API_TOKEN=...
```

Never commit or persist this token in DATA tables. Internal API calls are audited without storing the token or raw client IP address. Mutating administration actions are also audited; only the names of submitted fields are stored, never passwords or field values. Runtime cleanup applies retention to both audit tables.

Public and internal API rate limits, pagination limits, cache TTLs, audit retention and CORS origins are configurable in `.env.example`. In production, set at minimum `APP_ENV=production`, `APP_DEBUG=false`, a production `APP_URL`, HTTPS at the reverse proxy/web server, a strong internal API token, and production-only `DATA_API_ALLOWED_ORIGINS`.

The API is read-only in this milestone. It does not unlock provider staging, research execution, AI agents, canonical auto-apply or canonical population. Existing DATA construction locks remain unchanged.

## Integrity, private snapshots, external replicas and recovery drills

`v0.4.0-dev.14` introduced integrity checks and verifiable private database snapshots. `v0.4.0-dev.15` extends that safety layer with a second filesystem replica target and a real isolated restore drill, still before any real DATA population is allowed.

The integrity scanner is non-destructive. It checks the active database connection, pending migrations, required platform tables, public ULID integrity, a set of critical foreign-key relationships and the current construction-lock state. The latest report is stored privately under `storage/app/private` and can be refreshed from the administration page **Backups & integrity** or from the CLI:

```bat
php artisan cw:data:integrity
```

For a manual maintenance window, the deeper scan also checks extended historical/runtime relationships and duplicate public identifiers:

```bat
php artisan cw:data:integrity --deep
```

A daily standard integrity scan is enabled by default at `03:40`. It does not call providers, run AI agents or modify canonical DATA. The deeper scan is intentionally manual so future high-volume observation tables are not traversed every night.

Database snapshots are private and include a JSON manifest with the DATA platform version, creation time, database driver, schema fingerprint, file size, SHA-256 checksum, verification result and the state of every DATA construction lock at snapshot time. MySQL/MariaDB snapshots use `mysqldump` / `mariadb-dump` with a temporary protected client option file so the database password is not passed on the command line. Laragon installations are detected automatically when the binary is not already in `PATH`. SQLite snapshots use `VACUUM INTO` for a consistent copy.

Create a manual snapshot with:

```bat
php artisan cw:data:snapshot --label=manual
```

List or re-verify snapshots:

```bat
php artisan cw:data:snapshots
php artisan cw:data:snapshot:verify latest
```

Before a restoration drill, run the read-only preflight:

```bat
php artisan cw:data:restore:check latest
```

The preflight **never imports the dump and never modifies the current database**. It verifies the snapshot file, size, checksum, database structure and availability of the local restore client when applicable.

### External/off-machine snapshot replica

Set an absolute destination outside both the Laravel project and the primary snapshot directory:

```env
DATA_SNAPSHOT_REPLICA_PATH=\\nas\coaster-world-data-backups
```

On Windows, a UNC/NAS path is the preferred real off-machine destination. A second local disk or synchronized folder can also be used for testing, but it is not equivalent to a separate machine. DATA refuses project-local or primary-snapshot-overlapping destinations and never exposes the configured path through the internal API.

After the destination folder exists and is writable:

```bat
php artisan cw:data:snapshot:replicate latest
```

The copied database file is re-hashed after transfer. A private manifest is written beside it, the local manifest records only the replica state/type (not the destination path), and external retention is applied independently.

Automatic external replication remains disabled until the destination is validated:

```env
DATA_SNAPSHOT_REPLICA_SCHEDULE_ENABLED=false
DATA_SNAPSHOT_REPLICA_SCHEDULE=03:00
DATA_SNAPSHOT_REPLICA_RETENTION_DAYS=45
DATA_SNAPSHOT_REPLICA_RETENTION_COUNT=30
```

### Real isolated restore drill

After creating a fresh `dev.15` snapshot, run:

```bat
php artisan cw:data:restore:drill latest
```

For MySQL/MariaDB, DATA creates a generated temporary database on the same server, refuses SQL containing `USE`, database create/drop/alter statements or qualified references to the active DATA database, restores the dump into the temporary database, compares the restored table list with the snapshot schema fingerprint, verifies the migrations table when present, then drops the temporary database. For SQLite, the drill operates on a separate private file and deletes it afterwards. The active DATA database is never selected as the restore target.

The latest drill report is stored privately and shown on **Administration → Backups & integrity**. `--keep` exists only for deliberate manual inspection; normal validation should run without it so the isolated target is removed automatically.

Automatic database snapshots intentionally remain disabled during the construction phase:

```env
DATA_SNAPSHOT_ENABLED=true
DATA_SNAPSHOT_SCHEDULE_ENABLED=false
DATA_SNAPSHOT_SCHEDULE=02:30
DATA_SNAPSHOT_RETENTION_DAYS=30
DATA_SNAPSHOT_RETENTION_COUNT=20
DATA_SNAPSHOT_REPLICA_PATH=
DATA_SNAPSHOT_REPLICA_SCHEDULE_ENABLED=false
DATA_RESTORE_DRILL_TIMEOUT_SECONDS=1800
DATA_INTEGRITY_SCHEDULE_ENABLED=true
DATA_INTEGRITY_SCHEDULE=03:40
```

Snapshot files remain under the private Laravel storage disk by default (`storage/app/private/data-snapshots`) and are not exposed by any HTTP download route. External copies are also never exposed by DATA. The snapshot and replica schedules remain disabled by default until the manual replica and restore-drill validation has succeeded on the deployment host.

This milestone does not change or open any DATA population, provider staging, provider scheduling, canonical auto-apply, AI-agent or autonomous-research lock.

### v0.4.0-dev.15.1 restore drill fingerprint compatibility

The isolated restore drill normalizes schema-qualified table names before comparing the snapshot fingerprint. This keeps snapshots created by v0.4.0-dev.15 compatible when Laravel recorded names as `database.table` while the MySQL restore client returned plain `table` names. The correction only affects verification; it does not change snapshot contents or restore behavior.

### v0.4.0-dev.15.2 MySQL snapshot schema scope correction

Older `dev.15` snapshots may contain a legacy fingerprint generated from Laravel's unscoped MySQL table listing. On a development server hosting several databases, that listing can contain schema-qualified tables from databases other than Coaster World DATA even though `mysqldump` correctly exports only the active DATA database. `dev.15.2` scopes new fingerprints to the active MySQL database and, when validating a legacy snapshot, filters schema-qualified fingerprint entries to the database name stored in that snapshot manifest. This is a verification-only correction: the dump file, restore target, active DATA database and construction locks are not modified.

## v0.4.0-dev.16 operational monitoring and persistent alerts

This milestone adds a final operational health layer before any real DATA population is enabled.

Run the health evaluator manually with:

```bat
php artisan cw:data:health
```

The command evaluates database availability, construction locks, scheduler/worker heartbeats, queue backlog, failed jobs, provider circuit breakers and connectivity health, stale ingestion/AI runs, integrity freshness, snapshot freshness, external replica freshness, isolated restore-drill freshness, internal API configuration, recent API 5xx errors and free disk space.

The same evaluation is scheduled every five minutes by default and synchronizes persistent operational alerts in `system_alerts`. Alerts automatically resolve when their underlying check becomes healthy again. An administrator may acknowledge an active alert, but acknowledgement never hides or manually resolves the underlying problem. A warning that escalates to critical is reopened automatically.

The monitoring UI is available under **Administration → Monitoring & alerts**. The internal status API also exposes the live monitoring summary so an external watchdog can detect a dead scheduler even when the local scheduler can no longer run the periodic health command.

Default thresholds are intentionally conservative and can be overridden through the `DATA_HEALTH_*` environment variables documented in `.env.example`. During the construction phase, `DATA_HEALTH_EXPECT_LOCKS_CLOSED=true` makes any unexpected opening of a DATA safety lock a critical alert.

The monitoring scheduler and alert table are operational only: they never populate canonical DATA, never enable provider collection, never apply canonical decisions and never open a construction lock.


Apply the schema with the normal Laravel migration path:

```bat
php artisan migrate
```

A manual MySQL/MariaDB alternative is also supplied at `database/migrations/sql/coaster_world_data_v0.4.0-dev.16_system_alerts.sql`. Use one method only; do not apply both.


### v0.4.0-dev.16.1 monitoring Blade parse fix

This patch corrects the monitoring page translation call for `health_checks_hint`: the replacement-parameter array is now closed with `]` before the translation helper call is closed. The fix is presentation-only and does not change monitoring thresholds, alerts, schedules, migrations, DATA contents or safety locks.

## v0.4.0-dev.17 safe resilience and non-regression suite

This milestone adds a deliberately non-destructive resilience harness before any DATA population lock is opened.

Run the local resilience suite with:

```bat
php artisan cw:data:resilience
```

The command does **not** call external providers, does **not** stage provider data, does **not** run live AI research and does **not** write canonical DATA. It validates the construction locks and write guards, provider-job uniqueness, timeout/rate-limit/upstream failure classification, AI budget detection, simulation isolation, public/internal API limiter registration, byte-level snapshot corruption detection, stale-run recovery dry-run behavior and protected DATA counters before/after the suite.

The snapshot-corruption scenario runs only inside a generated private temporary directory and is deleted immediately afterwards. The latest report is stored privately at `storage/app/private/data-operations/resilience-latest.json` by default and its status is shown by `php artisan cw:data:status`.

Useful variants:

```bat
php artisan cw:data:resilience --json
php artisan cw:data:resilience --strict
php artisan cw:data:resilience --no-store
```

The deeper regression cases are covered by the PHPUnit resilience tests, including recovery of an interrupted ingestion run and a real HTTP burst against the public API limiter:

```bat
php artisan test --filter=Resilience
```

Existing canonical-engine, provider-normalization, AI runtime, API contract, snapshot/restore and monitoring tests remain part of the full regression suite:

```bat
php artisan test
```

The resilience milestone introduces no database migration and leaves all six construction locks unchanged.

## v0.4.0-dev.18 preproduction, clean rebuild and operations runbook

This milestone adds a final deployment rehearsal before the **Ready for DATA** audit. It still does not open any population/provider/AI/canonical lock.

A real clean-install rehearsal is available with:

```bat
php artisan cw:data:install:drill
```

The command creates a separate generated database target, applies every migration, runs the stable reference seeders, verifies required platform tables and reference-data minima, runs the seeders a second time to detect duplicate-producing seed logic, then removes the temporary target. The active DATA database is never migrated, seeded or selected by the drill.

The full readiness gate is:

```bat
php artisan cw:data:preflight
```

On the eventual production host use the stricter form:

```bat
php artisan cw:data:preflight --production --strict
```

The production mode additionally requires production environment hardening, MySQL/MariaDB, the database queue, healthy runtime heartbeats and enabled backup/replica/integrity/health schedules. Reports are stored privately and contain statuses only, never secrets.

Windows helpers are supplied for Laragon validation and recovery:

```bat
tools\verify_fresh_install_windows.bat
tools\preflight_data_windows.bat
tools\preflight_data_windows.bat production
tools\recover_data_runtime_windows.bat
tools\recover_data_runtime_windows.bat apply
```

The internal API now supports a temporary previous token during a coordinated rotation. Use `DATA_INTERNAL_API_PREVIOUS_TOKEN` only for the short grace window while clients move to the new `DATA_INTERNAL_API_TOKEN`, then remove it and clear Laravel's cached configuration. `cw:data:preflight` warns while a previous token is still configured.

The detailed reconstruction, runtime, crash-recovery, token-rotation and backup procedure is documented in `docs/PREPRODUCTION_RUNBOOK.md`.

## v0.4.0-dev.19.1 internal API test correction

This patch makes the internal API contract test independent from the persisted Ready-for-DATA audit state. After a successful `cw:data:ready:audit`, `data.operations.ready_for_pilot` is legitimately `true`; the API test now validates that the field remains a boolean instead of incorrectly forcing `false`. No runtime DATA logic, lock, schema, or pilot gate behavior changes.

## v0.4.0-dev.19 final Ready-for-DATA audit

This milestone is the final certification gate before the first real, deliberately small DATA pilot. It does not contact external providers and does not write synthetic data into the active DATA database.

Run:

```bat
php artisan cw:data:ready:audit
```

The command requires all six construction locks to be closed. It creates a disposable isolated database, applies migrations and stable reference seeders, then sends synthetic provider records through the complete platform path:

`provider registry -> normalization -> verified external mapping -> staging -> evidence-backed proposals -> consensus/conflict -> canonical decision -> isolated canonical write simulation -> public API resource`

The audit verifies both safe paths:

- two independent providers agreeing on the same value produce one explainable candidate;
- two independent providers disagreeing on a material value create a persistent conflict instead of silently overwriting canonical DATA.

Before the isolated canonical-write simulation, the normal write guard is explicitly tested and must reject the candidate while the construction lock is closed. Only inside the disposable audit database, the population/auto-apply flags are temporarily enabled in process memory to verify the actual canonical writer and public API projection. The flags are restored immediately afterwards, the isolated database is deleted, and active DATA row counts plus all six locks are compared before/after.

Expected output:

```text
Result: VERIFIED
Pilot gate: READY
```

The latest report is stored privately under `storage/app/private/data-operations/ready-for-data-latest.json` by default and is summarized by:

```bat
php artisan cw:data:status
php artisan cw:data:preflight
```

The first real pilot must remain manual. When the audit is `VERIFIED`, only `DATA_PROVIDER_COLLECTION_ENABLED` may be enabled initially, for explicitly named provider/park runs. Keep provider schedules, canonical population, canonical auto-apply, AI agents and autonomous research locked until staged records, mappings, proposals and conflicts from the pilot have been reviewed.

See `docs/READY_FOR_DATA.md` for the pilot gate procedure.

## v0.5.0-dev.1 first real provider pilot

The DATA foundation has passed the Ready-for-DATA gate, so this milestone introduces the first real provider contact against the active DATA server while keeping canonical writes impossible.

The initial pilot is intentionally narrow:

- provider: ThemeParks.wiki only;
- at most 3 explicitly selected parks;
- manual commands only;
- provider schedules remain locked;
- staging only: no proposals and no canonical catalogue writes.

Discover exact park IDs without writing:

```bat
php artisan cw:data:pilot:discover "Efteling"
```

Preview the exact target set:

```bat
php artisan cw:data:pilot:collect --park=<PARK_ID>
```

After a fresh verified and replicated snapshot, enable only:

```env
DATA_PROVIDER_COLLECTION_ENABLED=true
```

Then stage the reviewed target set:

```bat
php artisan optimize:clear
php artisan cw:data:pilot:collect --stage --park=<PARK_ID_1> --park=<PARK_ID_2>
php artisan cw:data:pilot:status --records
```

The pilot report proves that proposals, conflicts, applied decisions and canonical park/attraction/POI counts did not change. Existing identical staging records are preserved rather than re-linked to the pilot, which makes targeted cleanup safe:

```bat
php artisan cw:data:pilot:cleanup <PILOT_ID>
```

During this phase `cw:data:integrity`, `cw:data:health`, `cw:data:preflight` and `cw:data:status` recognize the exact safe state where only provider staging is enabled after a verified Ready-for-DATA audit. Any additional open DATA lock is still treated as unsafe.

See `docs/FIRST_DATA_PILOT.md` for the complete pilot runbook.

## Pilotage web complet et console sécurisée — v0.5.0-dev.2

L'administration devient le poste de pilotage principal de Coaster World DATA. Les opérations courantes du premier pilote ne nécessitent plus le terminal Windows :

- **Administration → Pilote DATA** gère recherche, preview, snapshot de sécurité, vérification/réplication, armement one-shot du staging, import staging, inspection et nettoyage ciblé ;
- l'autorisation web du staging est **désarmée par défaut** et doit être armée explicitement juste avant l'import ; elle est automatiquement retirée après une tentative de staging ;
- `DATA_PROVIDER_COLLECTION_ENABLED=true` reste uniquement un **kill-switch serveur maximal** : si cette permission est fermée, le web ne peut pas la contourner ;
- **Administration → Console DATA** fournit une console de maintenance avec liste blanche stricte. Elle exécute des commandes Artisan DATA approuvées sans passer par un shell système ;
- la console refuse notamment `migrate:fresh`, `tinker`, les opérateurs shell et toute commande hors liste blanche ;
- les sorties de la console sont limitées et conservées dans `storage/app/private/data-console/history.json` (chemin configurable) ;
- les actions restent admin-only, CSRF-protected, throttled et couvertes par l'audit administrateur existant.

Le terminal local reste utile pour installer un patch ou intervenir si l'interface web elle-même est indisponible, mais il n'est plus nécessaire pour piloter le workflow DATA courant.


### Correctif v0.5.0-dev.2.1

- l’armement web du staging est désormais strictement **one-shot** : toute tentative de staging depuis l’administration le désarme, y compris lorsqu’un prérequis (snapshot ou preview correspondante) bloque l’opération avant l’import ;
- l’intégrité, le monitoring, le preflight et `cw:data:status` distinguent maintenant correctement le **kill-switch serveur du pilote** de l’armement web temporaire : `DATA_PROVIDER_COLLECTION_ENABLED=true` reste sain tant que les cinq autres verrous sont fermés et que l’audit Ready-for-DATA est valide ;
- l’API interne expose séparément `pilot_master_staging_safe` et `pilot_staging_safe` (ce dernier exige toujours l’armement web actif).

## v0.6.0-dev.1 — administration simplifiée

L'administration est désormais centrée sur deux écrans destinés à l'usage quotidien :

- **Recherche & enrichissement** : saisir le nom d'un parc, coaster ou attraction, rechercher simultanément dans la base locale et les fournisseurs actifs, puis préparer un plan d'enrichissement providers-first avec IA en dernier recours ;
- **Fournisseurs de données** : gérer les adapters intégrés et ajouter des API JSON génériques comme des plugins configurables depuis le web.

Les pages historiques de pipeline, missions, monitoring, snapshots et console restent disponibles dans **Outils avancés** mais ne constituent plus le parcours normal.

Les plugins JSON peuvent définir une URL de recherche, un chemin vers la liste de résultats, les chemins ID/nom/type/URL, des capacités, des headers, un mapping de champs ainsi qu'une authentification Bearer/header/query. Les secrets sont stockés avec le cast chiffré Laravel et ne sont jamais réaffichés.


## Simple enrichment workflow (v0.6.0-dev.2)

The normal DATA workflow now stays on **Administration → Search & enrich**:

1. enter a park, coaster or attraction name;
2. select the matching result;
3. DATA queries active built-in providers and enabled JSON provider plugins;
4. provider values are merged field-by-field with source and confidence;
5. conflicting values remain visible instead of being silently overwritten;
6. an administrator selects the values to keep and clicks **Review and save**;
7. only these reviewed values are written to the canonical record, with field provenance.

For an external attraction/coaster the administrator chooses its parent park before enrichment. For a new external park DATA creates the canonical park only at the final review/save step. The selected external provider identifier is then stored as a verified mapping.

Wikipedia search results expose their Wikidata entity identifier when available, allowing the workbench to chain a structured Wikidata enrichment automatically. Geospatial sources can then be queried using the best coordinates already found. Enabled DATAtourisme/Foursquare providers and generic JSON plugins also participate automatically when configured.

AI remains a last-resort option. Missing fields are shown clearly in the same page; AI agents stay disabled unless the existing AI/research environment locks are explicitly enabled. The advanced Pipeline/Research pages remain available under **Advanced tools**, but they are no longer required for normal manual enrichment.
