graflo.hq.ingestion_parameters¶
Ingestion parameters and per-document cast-failure models for the caster.
This module exists to keep graflo/hq/caster.py focused on casting logic, while
keeping ingestion-policy types stable and importable.
Classes¶
CastBatchResult
¶
Bases: BaseModel
Outcome of casting a batch through a resource (possibly with skipped documents).
Source code in graflo/hq/ingestion_parameters.py
DocCastFailure
¶
Bases: BaseModel
Structured record for one source document that failed during resource casting.
Source code in graflo/hq/ingestion_parameters.py
Attributes¶
doc_index
instance-attribute
¶
doc_preview = Field(default=None, description='Subset or truncated JSON of the source document for debugging.')
class-attribute
instance-attribute
¶
exception_type
instance-attribute
¶
failure_kind = 'document'
class-attribute
instance-attribute
¶
location_path = Field(default=None, description='Extraction location path when failure_kind is transform.')
class-attribute
instance-attribute
¶
message
instance-attribute
¶
nulled_fields = Field(default=None, description='Output fields set to None when failure_kind is transform.')
class-attribute
instance-attribute
¶
resource_name
instance-attribute
¶
traceback = Field(default='', description='Formatted traceback, truncated to the configured max length.')
class-attribute
instance-attribute
¶
transform_label = Field(default=None, description='Transform name or module.foo when failure_kind is transform.')
class-attribute
instance-attribute
¶
DocErrorBudgetExceeded
¶
Bases: RuntimeError
Raised when total document cast failures exceed IngestionParams.max_doc_errors.
Source code in graflo/hq/ingestion_parameters.py
Attributes¶
doc_error_sink_path = doc_error_sink_path
instance-attribute
¶
limit = limit
instance-attribute
¶
total_failures = total_failures
instance-attribute
¶
Methods:¶
__init__(*, total_failures, limit, doc_error_sink_path)
¶
Source code in graflo/hq/ingestion_parameters.py
IngestionParams
¶
Bases: BaseModel
Parameters for controlling the ingestion process.
max_items caps how many source items (rows, JSON objects, grouped
RDF subjects, …) are read per resource run. It maps to
AbstractDataSource.iter_batches(..., limit=...). batch_size is only
the maximum number of items per yielded batch, not a cap on total volume.
Source code in graflo/hq/ingestion_parameters.py
76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | |