Build your own: the repository, three commands, and what to do next

This lesson enables the reproduction and extension of a structured approach to orbital vocabulary design using open-source tools and data. The method is built around a falsification harness, which allows for testing and refining vocabularies through practical application. The process is designed to be repeatable and transparent, with results that can be regenerated using a defined set of commands.
Reproduction: Three commands and a pinned snapshot
Reproduction of the course’s results requires executing three commands as documented in the repository’s README. These commands build the graph from a pinned snapshot, run the falsification harness, and execute the rules. The entire process takes minutes to complete on a standard laptop. The method uses two open-source libraries and is fully reproducible, with every number in the course regenerated from the same inputs. The pinning ensures consistency across runs and allows for meaningful comparisons.
Extension one: Refreshing the catalogue
The first extension involves refreshing the catalogue with today’s SATCAT data, re-pinning the snapshot, and re-running the rules. The resulting differences between the old and new outputs are significant, as they highlight launches, decays, and status changes. The pinning discipline ensures that these deltas are meaningful and not artifacts of data drift. This process is a practical demonstration of how the system evolves over time with new data.
Extension two: Cross-catalogue integration
The second extension introduces DISCOS objects and treats identity resolution as an alignment problem with extensional evidence. This approach leverages the defence-relevant exercise of lesson 13, where the goal is to align vocabularies across different catalogues. The method allows for cross-catalogue comparisons and provides a framework for handling discrepancies in object identification. The system treats this as a validation task, ensuring that entities are correctly matched across datasets.
Extension three: Matcher evaluation
The third extension evaluates any ontology matcher over the lifted-vocabulary-to-SSAO (Space Situational Awareness Ontology) pair. The system scores the full candidate set of matches through both channels, providing a comprehensive evaluation framework. This harness accepts arbitrary correspondence lists, which allows for flexible testing of different matching strategies. The results can be published within an afternoon, making it a practical approach for rapid iteration and validation.
Extension four: The gate in anger
The fourth extension involves deploying an LLM (Large Language Model) extractor over real launch announcements. The output of the LLM is then staged and measured against both closed-world and extensional gates. This process identifies what information the system captures, forming a fine-tuning curriculum for the LLM. The captured set of outputs provides a realistic dataset for improving model performance in real-world scenarios.
Where this fits in the curriculum
This lesson is part of a broader curriculum that spans defence, industrial, construction, and orbital vocabularies. It is supported by five arXiv papers, public repositories, and a companion deep-dive paper currently in preparation. Tesseract Academy teaches and consults across the stack, with linked research pages carrying a summary article for this work. The approach taken here is designed to be both academically rigorous and practically applicable.
What to take away
The course provides a framework for building and testing orbital vocabularies using reproducible steps and open tools. The extensions offer clear paths for evolving the system with new data, integrating different catalogues, and evaluating matching strategies. The final extension demonstrates how to apply the method in real-world settings using LLMs and practical gates. These techniques can be applied to similar problems in other domains with minimal modification.
Reference
| Lesson | 15 of 15 |
| Outcome | Reproduce the course’s results and extend them in four concrete directions. |
| Worked repository | neurosymbolic-space-kg on GitHub |
| Domain ontology | Space Situational Awareness Ontology (Rovetto) |
Sources and further reading
Keep going with Tesseract Academy
Get personalised feedback from the Tesseract Academy AI Coach
Paste your work below and the AI Coach will review it against a rubric and suggest concrete improvements.
Two tools to take this further:
- Tesseract Academy AI Coach – a personal AI coach that reviews your work, runs practice scenarios and helps you apply what you learned.
- LearnGen – build your own AI-powered courses and training with Tesseract Academy’s course engine.
