All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
QuickMapper—unit_from_file: trueon a column annotation: reads the unit string from the file's units row (the row thatskip_after_headerskips) and resolves it to a QUDT IRI automatically. Resolved and unresolved results are included inresult.oold_doc["unit_resolutions"].QuickMapper: columnunit:plain strings (e.g."mm","N","s","%") are now automatically resolved to QUDT IRIs using the same built-in alias table. Full IRIs are passed through unchanged.QuickMapper: column descriptors are now created for every annotated column — not only columns that have anirikey. A column with only aunitannotation still gets acsvw:Columnnode in the graph with anobo:IAO_0000039unit triple.QuickMapper: metadataunit:plain strings are auto-resolved to QUDT IRIs before the triple is written (e.g.unit: "°C"→obo:IAO_0000039 uqudt:DEG_C).
Transformer._add_timeseries_nodes: time-series descriptor pattern (container_type,column_predicate,column_type,unit_predicate, …) is now read from the schema'stimeseries_patternblock rather than being hardcoded. Schemas that add atimeseries_patternkey control how the library serialises column descriptors; the default pattern usescsvw:Table/csvw:Column/obo:IAO_0000039— matching the TTO/PMDCo3 S355 reference dataset.QuickMapper: column descriptor pattern (column_predicate,column_type,column_name_predicate,unit_predicate) is now read from the mapping config's optionalcolumn_patternkey; defaults to the same csvw /obo:IAO_0000039pattern as the Transformer. The previous default (dcat:Datasetroot type,qudt:hasUnitunit predicate) is no longer emitted.QuickMapper:result.oold_doc["unit_resolutions"]andresult.oold_doc["metadata"][…]["unit"]now store the resolved QUDT IRI whenunit_column: trueresolves successfully, rather than the raw unit string from the file.ParseResult.column_iristype widened fromdict[str, str]todict[str, str | None]. Parsers that map a column to an ontology class may now emitNonefor columns that have no applicable class.unit_column: truerenamed tounit_from_file: truein mapping configs (both metadata fields and, new in this release, column annotations). The old key name is no longer recognised; update existing YAML configs accordingly.- TestXpertIII parser (
testxpert_iii/parser.py):column_irisnow emitsNoneforPrüfzeitandStandardkraft(no TTO v3 class); full TTO v3.0.0 numeric IRIs forDehnung(TTO_0000004),Standardweg(TTO_0000005),Breitenänderung(TTO_0000011),Traversenweg absolut(TTO_0000013).
column_mapping.json(testXpert III / de):Dehnungunit corrected fromMilliMtoPERCENT— the example file's units row reports%for this column, and TTOTTO_0000004(elongation) is a dimensionless ratio expressed in percent. All TTO class references updated from v2 human-readable IRIs to TTO v3.0.0 numeric IRIs ornullwhere no v3 class exists.- QuickMapper quickstart notebook (
docs/3_quickstart-mapping.ipynb): column annotations updated to TTO v3.0.0 numeric IRIs; plain-string units used throughout to demonstrate automatic unit resolution; root type updated fromdcat:Datasettocsvw:Tablein the description cell.
Consumers reading result.column_iris should guard against None values:
iri = result.column_iris.get(col)
cls = iri.rsplit('/', 1)[-1] if iri else '—'Consumers reading result.oold_doc["metadata"][field]["unit"] for unit_column: true
fields will now receive a QUDT IRI string instead of the raw unit string.
Mapping configs that relied on dcat:Dataset as the default root type should set
root_type: "http://www.w3.org/ns/csvw#Table" explicitly (or omit root_type to
accept the new default).
- Notebook cells now print only filenames (
.name) instead of full absolute paths, so committed outputs do not expose the local machine's directory tree on GitHub.
unit_column: truenow automatically resolves unit strings to QUDT IRIs using a built-in lookup table covering common lab units (N, kN, mm, MPa, °C, s, %, and ~30 more). Resolved units are stored withqudt:hasUnit <IRI>; unrecognised strings fall back toqudt:unit "string"as before.result.oold_doc["unit_resolutions"]: dict mapping each file unit string to its resolved QUDT IRI, ornullif no match was found.QuickMapper.run()prints a resolution summary attributed to the source file (e.g.QuickMapper: unit resolution for 'my_file.TXT': ...) wheneverunit_column: truefields are present.
QuickMapperquickstart notebook: added a "What does a measurement file look like?" section; added unit resolution and unrecognised-unit example cells.- README folder structure updated to reflect the
parser.py+ locale-subfolder layout introduced in v0.2.0. docs/1_getting-started.md: terminology aligned (skip_rowscomment now says "column names row").
ZwickParserrenamed toTestXpertIIIParser; import path changes fromsemantic_transformers.parsers.characterization.tensile_test.zwicktosemantic_transformers.parsers.characterization.tensile_test.testxpert_iii. Update any existing imports.QuickMappermetadata field config: thepredicatekey is renamed toproperty. Update any existing YAML or dict configs that usepredicate:.
scripts/run_notebooks.sh— single entry point to run the test suite, validate notebooks, or refresh notebook outputs in-place.
QuickMapperquickstart notebook revised: added a conceptual "what does a measurement file look like?" section; clarifiedskip_rows,skip_after_header, andunit_column; removed internal variable names (_cwd,_candidates) from the file-path cell; removed em-dash constructions throughout.CONTRIBUTING.mdanddocs/1_getting-started.mdupdated to referencescripts/run_notebooks.shfor running tests and refreshing notebooks.
gauge_length/gauge_length_unit—Messlänge Standardwegmetadata row now parsed and emitted as apmdco:PMD_0000013process condition.preload/preload_unit—Vorkraftmetadata row parsed as a pre-load condition.test_date—Datum/UhrzeitExcel serial-number date auto-converted to an ISO 8601 datetime string via the new_excel_serial_to_iso()helper.unit_field_mapparameter — generalises unit-column extraction for any metadata label; replaces the hardcodedstrain_rate_labelmechanism.strain_rate_labelis retained for backwards compatibility but is deprecated.
flat_graphproperty — returns ardflib.Graphwith all triples and namespace bindings propagated from the internalDataset. Replaces the repetitivefor s, p, o, _ in result.graph.quads(): flat.add(...)pattern in notebooks.
- Namespace bindings in serialised TTL output (
pmdco,tto,obo,qudt, …) are now derived from the schema@contextand rdflib's built-in namespace manager rather than being hard-coded in Python. Adding a prefix to the schema YAML is sufficient; no library changes are needed.
ZwickParseris now compatible withcharacterization/tensile-test/TTOv1.1.0. Remains backwards-compatible with v1.0.0 files (all new fields are optional).
basekeyword-only parameter onTransformer.run()— pass a custom base IRI (e.g."https://example.org/") to override the schema's@baseentry so all data node IRIs are resolved against your own namespace instead of the schema's internal PMDCo test namespace
- Moved
parsers/from repo root intosrc/semantic_transformers/parsers/so parsers are included in the wheel and importable after a regularpip install - Renamed
tensile-test/totensile_test/throughout to produce valid Python package identifiers - Removed the
importlib-based shim insrc/semantic_transformers/parsers/that only worked with editable installs - Only
*.jsondata files (column mappings) are shipped in the wheel; parserREADME.mdfiles are excluded viapackage-data - Removed duplicate
example_tensile_test.TXTfrom the parser folder (canonical copy is intests/data/) - Cleaned up test imports to use the proper package path instead of
sys.pathmanipulation
- Example tensile test data bundled with ZwickParser
example_tensile_test.TXTnow included in parsers distribution- Users can load example files directly from installed package
- Parsers module now included in PyPI distribution
- Zwick/Roell tensile test parser available via
from semantic_transformers.parsers import ZwickParser - Users can now access sample parsers without requiring GitHub checkout
- Zwick/Roell tensile test parser available via
- Initial public release of semantic-transformers
- Core library components:
Transformer: Main class for running parsing → JSONata transform → RDF graph pipelineParser: Protocol for implementing custom instrument parsersParseResult: Standard return type for parsers (simplified JSON + DataFrame)TransformResult: Pipeline result containing RDF graph and metadataQuickMapper: Simple YAML-based mapping system for tabular files without custom parsers
- Characterization parsers:
- Tensile test parser for Zwick/Roell (testXpert III) instruments
- Support for both CSV and binary data formats
- Multi-format file support: CSV, TSV, Excel (.xlsx), Parquet, JSON
- Automatic RDF/Turtle graph generation
- DataFrame export for data inspection
- JSONata transformation templates
- Column mapping to ontology IRIs and units
- Optional dependencies for Excel file handling
- Getting Started Guide (docs/1_getting-started.md)
- Parser Development Guide (docs/2_adding-a-parser.md)
- QuickMapper Quickstart Notebook (docs/3_quickstart-mapping.ipynb)
- Example measurement data (docs/example_measurement.ttl)
- Full API documentation and examples in docstrings
- Comprehensive test suite for Transformer, Parser, and QuickMapper classes
- Example data and test fixtures included
- Tests for column mapping and unit conversion
- pandas
- rdflib
- pyyaml
- jsonata-python
- jsonschema
excel: openpyxl (for Excel file support)dev: pytest, nbmake (for development and testing)
- Python 3.10+
- testxpert_iii schema compatibility table updated to reflect the
semantic-schemasv0.6.0 release: schema versions reset to0.1.0(SemVer 0.x pre-release convention), andcharacterization/tensile-test/PMDCoadded as a second compatible schema alongside TTO.
- Additional instrument parsers (metallography, microscopy, analysis)
- Streaming support for large data files
- Caching and memoization for repeated transformations
- Web API wrapper for parser services
- Expanded unit conversion and validation
- Database output formats (JSON-LD, RDF-JSON)