WIKAI: A Living Dictionary (GALACTICA) and Kolmogorov Ontology (ALGORITMICA)
★ Giulio Ruffini, Francesca Castaldo,
★ guarantor: Giulio Ruffini · vouches for the paper per WP0084 §6
WIKAI is a two-component research program that addresses the problem of semantic drift and conceptual fragmentation in scientific and clinical language by combining continuous definition regeneration with compression-driven ontology induction. The first component, WIKAI GALACTICA, maintains a living dictionary whose entries are periodically re-proposed by a retrieval-augmented large language model operating over a governed, provenance-tracked corpus, optimizing a multi-objective function that balances coverage, parsimony, entailment fidelity, and readability under a structured genus–differentia schema. The second component, WIKAI ALGORITMICA, applies a Kolmogorov-complexity-inspired minimum description length framework to induce and refactor concept graphs, proposing machine-coined neologisms when a latent regularity in the literature lacks a precise lexical handle and a new term would reduce total description length. Both components share provenance-first retrieval, diachronic embedding indices for sense-drift detection, and a governance layer that enforces curator oversight, freeze policies for sensitive clinical terms, and anti-recursion safeguards against model-collapse from AI-generated text. The framework is applicable to general or domain-specific corpora and is evaluated both intrinsically—through expert adequacy ratings and NLI-based inclusion/exclusion probes—and extrinsically through gains in word-sense disambiguation, entity linking, and clinical note disambiguation.
A two-part system that keeps scientific vocabulary alive: one half continuously rewrites definitions from fresh literature, the other invents new terms when the existing vocabulary is provably inadequate.
Language in science doesn't stand still. Words like "depression" or "consciousness" mean different things to a clinician, a cognitive scientist, and a physicist — and those meanings shift as research accumulates. Standard dictionaries freeze at publication. Wikipedia is ungoverned. LLMs hallucinate. WIKAI is BCOM's answer: a governed, retrieval-grounded system that treats vocabulary maintenance as an ongoing optimization problem rather than a one-time editorial act.
The first half, GALACTICA, is the near-term deliverable. It runs a retrieval-augmented language model over a curated, continuously updated corpus — journals, preprints, clinical guidelines — and periodically re-proposes dictionary entries in a strict genus–differentia format (what category does this thing belong to, and what distinguishes it from everything else in that category). Each definition comes with inclusions, exclusions, citations, usage examples, and register notes (so the same term is handled correctly in a clinic versus a paper versus a classroom). Crucially, when usage measurably drifts — tracked via embedding-space divergence across time windows — the system proposes a revision with side-by-side redlines for human curators to approve. The worked examples in the paper (for "depression," "consciousness," "life") show concretely how much richer and more precise these entries are compared to a standard dictionary gloss.
The second half, ALGORITMICA, is the longer research horizon and the more ambitious idea. Instead of asking "what's the best definition of a term people already use," it asks "what conceptual structure best compresses the literature we actually have?" It maintains an explicit typed graph of concepts and relations, and periodically searches for edits — merges, splits, new relations — that reduce the total description length of the corpus (a Kolmogorov/MDL framing: shorter is better if it predicts just as well). When a recurring pattern in the data has no clean name, ALGORITMICA can propose one — a neologism — with a formal back-definition anchored in existing terms. These coined terms live in a sandbox and are never forced into use; GALACTICA tests uptake before any promotion.
Both halves share the same infrastructure commitments: provenance-first (every definition token traces to a source), anti-recursion filtering (AI-generated text is aggressively excluded from the training corpus to prevent model collapse), multilingual tracking, and human curator oversight for sensitive clinical terms. The distinction between GALACTICA's human-established vocabulary and ALGORITMICA's machine-coined constructs is kept explicit and visible to users at all times — a meaningful integrity guarantee.
The practical upside is cleaner manuscripts, faster cross-lab onboarding, better entity linking, and reduced ambiguity in clinical notes. The strategic upside is ALGORITMICA's capacity to discover and name conceptual structure that the field hasn't yet articulated — a principled alternative to the usual process of terminology emerging haphazardly from whoever publishes first.
- Zenodo
- 10.5281/zenodo.21008526
- WP ID
- WP0038
- Lifecycle
- ongoing
- Visibility
- internal
- Access level
- open
- Embargo until
- —
- Priority
- —
- Collab
- closed
- Venue
- —
- DOI
- —
- Deadline
- —
- Owner
- —
- Source
- drive_legacy
- Repo path
- WP0038 - WIKAI GALATICA
- v0.1.0 (draft) · drive-legacy · zenodo:21008527Auto-created by Phase 1a bootstrap ingestion.
