Download the PHP package amateescu/prov without Composer
On this page you can find all versions of the php package amateescu/prov. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download amateescu/prov
More information about amateescu/prov
Files in amateescu/prov
Package prov
Short Description PHP implementation of the W3C Provenance Data Model (PROV-DM)
License MIT
Informations about the package prov
prov: W3C Provenance for PHP
PHP implementation of the W3C Provenance Data Model (PROV-DM).
PROV-DM describes where things come from: entities (things you care about), activities (things that happen), and agents (who's responsible). Relations like wasGeneratedBy and wasAttributedTo connect them to form a provenance graph.
PROV-DM fits data lineage, audit trails, scientific-workflow provenance, attribution graphs, and any case where you need to record where information came from.
This library provides a fluent builder for assembling that graph, round-trip serializers for PROV-JSON, PROV-N, and PROV-XML (plus serialize-only PROV-JSONLD), document operations (merge, flatten, semantic equality), and a partial PROV-CONSTRAINTS validator.
Requirements
- PHP 8.4+
ext-dom(only if you useXmlSerializer)
Installation
Quick start
The static Prov:: calls are a convenience facade. Under a dependency-injection container, construct the underlying classes directly: DocumentBuilder, the per-format serializers (Format::Json->createSerializer() / createDeserializer()), and ConstraintValidator.
Always pass relation arguments by name. PROV-DM fixes a per-relation positional order that does not follow subject-before-object.
wasGeneratedBytakes(entity, activity)butusedtakes(activity, entity): the two sit in opposite orders even though they connect the same two records. Positional calls silently invert the relation:The optional relation identifier is the last parameter of every relation method, so a positional call binds endpoints, never the id. Pass it by name:
wasGeneratedBy(entity: ..., activity: ..., identifier: 'ex:gen1').
Format support
| Format | Serialize | Deserialize |
|---|---|---|
| PROV-JSON | yes | yes |
| PROV-N | yes | yes |
| PROV-XML | yes | yes |
| PROV-JSONLD | yes | no (would require an RDF-aware parser) |
Output ordering
PROV serializations are unordered (a document is a set of records, namespaces a set of declarations), so ordering never affects meaning. For stable, readable output every serializer always sorts namespace declarations: the prov/xsd built-ins first, then the rest alphabetically by prefix. Records keep the order you added them by default; pass sortRecords: true to a serializer to order them into PROV-DM concept order instead (elements first, then relations in component order, each group sorted by identifier):
Reserved prefixes
Every serializer binds prov and xsd itself, because it writes prov:* and xsd:* terms of its own: PROV-XML on the document root, PROV-JSONLD in the @context, PROV-N through the grammar. A document that binds either prefix to another namespace keeps the namespace but not the prefix, and the names under it are written through a minted prefix, so every name still expands to the URI the model carries.
The XSD namespace has two spellings, http://www.w3.org/2001/XMLSchema (what PROV-XML binds, so that xsi:type="xsd:int" names an XML Schema type) and the same URI with a trailing # (what PROV-JSONLD and PROV-N bind). They build a different URI for every name, so a document binding xsd to the spelling a format does not use gets a minted prefix there too. Literal datatypes are the exception: an XSD datatype is written against the format's own xsd binding whichever spelling the model carries, and the two spellings compare equal.
PROV-N notes
The PROV-N parser accepts two convenience extensions beyond the published grammar, so input that parses here is not necessarily canonical PROV-N: line (//) and block (/* */) comments, and optional commas between a relation's arguments. Output always uses the canonical form.
PROV-N has no slot for an explicit identifier on specializationOf, alternateOf, hadMember, or mentionOf, and hadDictionaryMember has room for neither an identifier nor attributes. When a document carries one of these relations with an identifier (legal in PROV-JSON/PROV-XML), the PROV-N serializer drops the identifier, since the grammar cannot express it. DocumentComparator::equals() will flag the difference on a JSON-to-PROV-N-to-JSON round trip; keep such relations in PROV-JSON or PROV-XML if their identifiers matter.
The prov and xsd prefixes are implicit in PROV-N and the grammar forbids redeclaring them, so the serializer never writes those declarations; a document that binds either prefix elsewhere is handled as described under "Reserved prefixes" above. Every other prefix is checked against the PN_PREFIX production before it is written, Unicode included, so a prefix that starts with a digit or ends in a dot is refused rather than emitted as unparseable text.
PROV-JSONLD notes
PROV-O models specializationOf, alternateOf, hadMember, and mentionOf as plain object properties, and PROV-Dictionary does the same for hadDictionaryMember. A statement in one of those forms is a single triple on the subject node, with no node of its own, so there is nowhere to write a relation identifier or extra attributes. Serializing a document whose record carries either throws Prov\Exception\ProvException rather than writing the triple without them. Every other relation has a qualified form (prov:qualifiedGeneration, prov:qualifiedInsertion, ...) that carries both.
A relation is written as a property of its subject node, so a relation whose subject formal is missing has no node to hang on. One that carries an identifier, attributes, or another formal throws Prov\Exception\ProvException; one that carries nothing else states nothing and is dropped. A relation that has its subject but no object formal keeps its qualified form, which PROV-O allows without the object, so nothing is lost. The five relations above have no qualified form, so a missing object throws there too.
The whole document shares one @context. Bundle declarations are promoted into it when their prefix is free there; a bundle prefix that rebinds a document prefix cannot be, and names under it are written with a minted prefix instead, so every compact IRI expands to the URI the model carries.
Document operations
Querying a document
ProvGraph indexes a document (or bundle) once and answers edge queries by identifier, accepting QualifiedName objects, prefix:local shorthands, or full URIs:
agentsOf() returns one AgentInvolvement per association, each carrying the agent, the plan the association named, the association's attributes (so prov:role survives), and onBehalfOf: the chain of agents this one acted on behalf of, walked from the actedOnBehalfOf delegations (nearest responsible first, activity-scoped delegations honored, cycles guarded). It reports identifiers and structure only; read an agent's prov:type with recordByIdentifier() to classify it.
The graph covers the container's own records; flatten a document first to query across bundle boundaries. For type-centric queries (all Usage records), Document::getRecordsByType() remains the right tool.
Scanning stored PROV-JSON
JsonScanner answers slice queries straight off decoded PROV-JSON, without building the Document object graph, for the case where a stored document is large and a read wants a few facts from it. Identifiers and attribute names match by URI, so the prefix a document spelled them with does not matter:
attributeValue() returns the value as decoded, so a dateTime, a language-tagged literal or a qualified-name reference comes back as its raw typed map; the typed reads and the bags apply the deserializer's mapping instead. Record-level damage is skipped, structural damage throws at construction. The scanner reads the top-level document only; deserialize to work across bundles.
Validation
Coverage is partial: rules that need transitive graph reasoning over derivation chains aren't implemented, so $isValid === true only means no checked rule was violated. Use ConstraintValidator::implementedConstraints() or ::unsupportedConstraints() to see the exact set.
Builder tips
Namespaces. Register namespaces one at a time (namespace(), addNamespace()) or in bulk from an application-wide registry (addNamespaces($iterable)); DocumentBuilder also accepts an iterable to preload at construction. Re-registering a prefix with a different URI throws, including the prov/xsd built-ins, so a typo cannot silently corrupt a binding. build() prunes the declarations down to the namespaces your records actually reference, so registering many namespaces up front does not bloat the serialized output; call keepUnusedNamespaces() to keep them all. Documents obtained from Prov::deserialize() are not affected: they keep every namespace they declared.
Attributes. Pass attributes as an associative array: keys are resolved as namespace shorthands, and a list value adds one entry per element (that is how a repeated key is written, since PHP array keys are unique):
String values stay string literals, with one exception: a prov:type value written as a registered shorthand ('prov:type' => 'ex:Document') resolves to a qualified name, because prov:type values name types rather than carry text. For every other key, a string like 'workspace:stage' is stored verbatim; pass a QualifiedName object when you mean a reference. Prov\Attribute\AttributesBuilder offers the same rules imperatively, useful when attributes accumulate across code paths:
Two Attributes bags combine with $a->merge($b): a multimap union that keeps all values under a shared key (the way to promote a single value to several).
Blank nodes (anonymous records):
Use QualifiedName::blankNode('b1') instead when you control the label.
Bundles. withBundle() is the recommended form: it builds the bundle eagerly, inline, without breaking the fluent chain:
Two alternatives exist for other flows: bundle() returns a detached BundleBuilder that you drive directly and that is built lazily when the document's build() runs, and addBundle() attaches an already-built Bundle (for example one obtained by deserializing).
DocumentBuilder::build() and BundleBuilder::build() are single-use; a second call throws LogicException.
Learn more
Every public class carries an inline docblock explaining what it's for. The most useful starting points:
Prov\Prov: the facade used in the examples aboveProv\Builder\DocumentBuilder: the full set of record and relation methodsProv\Format: supported serialization formatsProv\Constraint\ConstraintValidator: what each PROV-CONSTRAINTS rule checks
Development
Before submitting a PR, run composer check (format, lint, analyze, tests).
See also
trungdong/prov: Python implementation of PROV-DM.lucmoreau/ProvToolbox: Java toolkit for PROV.
License
This library is made available under the MIT License. Please see LICENSE for more information.
All versions of prov with dependencies
ext-json Version *