Translate source data without losing its meaning
Build an adapter that preserves identifiers, validates an explicit source contract and returns stable domain types. Test what can stay unchanged when a file becomes an API, and what requires a new business decision.
30 MIN
By FDEInterviews · Updated
TL;DR: Keep source-specific translation behind an adapter with an explicit contract. Preserve identity and meaning, reject unsupported values, and test the downstream behavior you expect to remain stable.
Where you are. You can read an unfamiliar system and locate its boundaries. Now you will write the boundary between a customer's order feed and a small reporting workflow, including the tests that reveal a destructive conversion.
A rename is easy; a changed meaning needs a decision
Consider a fictional freight integration. A daily export uses ORDER_NBR, cust_id and sts. The report reads these names directly, a totals function converts quantities, and the export job has its own interpretation of status 9. A second region introduces 09. You now have several places to investigate before deciding whether the records mean the same thing.
An adapter collects those source assumptions and produces types the rest of your application understands. If the source model and your model have different semantics, this translation boundary is often called an anti-corruption layer. It can live inside your application; it does not require a separate service. Microsoft's pattern guidance also discusses validation, consistency and the cost of maintaining the layer.
The boundary does more than rename fields. A customer identifier is an identifier even when it contains only digits. Removing its leading zeros can merge two customers. Converting the text to a number discards exactly the characters that told the two apart, and once they have collapsed no downstream code can recover which record was meant. A timestamp without a timezone does not become UTC because that would be convenient. And closed may mean financially closed in one system but physically delivered in another.
An adapter cannot settle those questions on its own. Its job includes making the agreed answers visible.
Keep source assumptions together
The diagram shows a small integration. A larger one may have several adapters for separate source contracts. Tests, fixtures and transport clients will also mention source field names. Counting files is a useful search technique, but there is no rule that one changed file proves good architecture and three prove bad architecture.
Use a more concrete review question: when a source representation changes without changing its meaning, can you update the translation and its tests without rewriting the reporting calculation? When meaning changes, which domain rules should change with it?
Write the contract before the conversion
The following contract and data are synthetic training material, not a customer's documented API. In a real engagement, record the source owner's confirmation and the contract version; do not invent either.
| Field | Contract for this exercise | What the adapter must avoid |
|---|---|---|
ORDER_NBR | Nonempty text; surrounding ASCII spaces are transport padding | Removing interior characters or accepting an all-space ID |
cust_id | Nonempty text, preserved exactly; no surrounding whitespace | Converting to an integer or stripping zeros |
order_dt | Ten-character YYYY-MM-DD business date | Guessing a timezone or slicing an arbitrary timestamp |
sts | 1 open, 9 and 09 closed, X cancelled | Inferring that every zero-padded code is equivalent |
ln_items | Nonempty list; each item has a SKU and positive integer quantity encoded as ASCII digits | Accepting booleans, floats, negative quantities or blank SKUs |
| Extra fields | Ignored by this report | Claiming every additive change is safe for every consumer |
The workflow needs total item quantity, not line-level prices or inventory reservations. That is why the domain type below contains a total. A fulfillment workflow would probably need the individual lines and their units; this type would be insufficient.
Save this complete example as adapter.py. It uses only Python's standard library.
from dataclasses import dataclass
from datetime import date
from enum import Enum
import re
class ContractError(ValueError):
pass
class OrderStatus(Enum):
OPEN = "open"
CLOSED = "closed"
CANCELLED = "cancelled"
@dataclass(frozen=True)
class Order:
number: str
customer_id: str
placed_on: date
status: OrderStatus
quantity: int
STATUS = {"1": OrderStatus.OPEN, "9": OrderStatus.CLOSED,
"09": OrderStatus.CLOSED, "X": OrderStatus.CANCELLED}
def text_field(raw, key):
value = raw.get(key)
if not isinstance(value, str) or not value:
raise ContractError(f"{key}: expected nonempty text")
return value
def to_order(raw):
if not isinstance(raw, dict):
raise ContractError("order: expected object")
number = text_field(raw, "ORDER_NBR").strip(" ")
if not number or number != number.strip():
raise ContractError("ORDER_NBR: invalid padding")
customer = text_field(raw, "cust_id")
if customer != customer.strip():
raise ContractError("cust_id: surrounding whitespace")
day = text_field(raw, "order_dt")
if not re.fullmatch(r"[0-9]{4}-[0-9]{2}-[0-9]{2}", day):
raise ContractError("order_dt: expected business date")
try:
placed_on = date.fromisoformat(day)
except ValueError as error:
raise ContractError("order_dt: invalid date") from error
status = STATUS.get(text_field(raw, "sts"))
if status is None:
raise ContractError("sts: unsupported status")
lines = raw.get("ln_items")
if not isinstance(lines, list) or not lines:
raise ContractError("ln_items: expected nonempty list")
total = 0
for line in lines:
if not isinstance(line, dict):
raise ContractError("line: expected object")
sku = text_field(line, "sku")
if not sku.strip():
raise ContractError("sku: blank")
qty = text_field(line, "qty")
if not re.fullmatch(r"[0-9]+", qty) or int(qty) <= 0:
raise ContractError("qty: expected positive integer text")
total += int(qty)
return Order(number, customer, placed_on, status, total)
SAMPLE = {
"ORDER_NBR": " A-4471 ", "cust_id": "00099421",
"order_dt": "2026-03-14", "sts": "09",
"ln_items": [{"sku": "X1", "qty": "2"},
{"sku": "X2", "qty": "3"}],
}
if __name__ == "__main__":
order = to_order(SAMPLE)
print(order.number, order.customer_id, order.status.value, order.quantity)
Run python adapter.py. Expect A-4471 00099421 closed 5. The zeros survive. The mapping for 09 is explicit. The date is validated as a date rather than extracted from whatever text happened to arrive.
This is a teaching adapter for a bounded in-memory object. A deployed transport also needs response-size limits, authentication, pagination, deadlines and protected diagnostics. The raised errors name fields without copying customer values into logs. The caller decides whether a rejected record can be held individually or whether the whole report must wait for correction.
Test the conversion that looks harmless
Run this second block beside adapter.py:
from copy import deepcopy
from adapter import SAMPLE, ContractError, to_order
first = deepcopy(SAMPLE)
second = deepcopy(SAMPLE)
second["cust_id"] = "99421"
assert to_order(first).customer_id != to_order(second).customer_id
assert len({first["cust_id"].lstrip("0"),
second["cust_id"].lstrip("0")}) == 1
for field, bad in [("sts", "009"), ("order_dt", "2026-02-30"),
("order_dt", "2026-03-14T00:00:00"),
("cust_id", " 00099421")]:
raw = deepcopy(SAMPLE)
raw[field] = bad
try:
to_order(raw)
except ContractError:
pass
else:
raise AssertionError(f"accepted invalid {field}")
raw = deepcopy(SAMPLE)
raw["ln_items"][0]["qty"] = True
try:
to_order(raw)
except ContractError:
pass
else:
raise AssertionError("accepted a boolean quantity")
print("identity preserved; five invalid cases rejected")
The first pair is the useful surprise: both source IDs contain digits, but numeric-looking text does not establish numeric identity. The deliberately destructive lstrip expression collapses them to one value. That is a correctness failure even if every downstream sum still runs.
Develop offline without taking data you cannot keep
Use synthetic fixtures like SAMPLE whenever they cover the behavior. If an approved captured response is needed to reproduce a defect, establish where it may be stored, who may access it, how it is transformed and when it must be deleted. Removing names alone may leave identifying IDs, free text or linked records. Permission to call an API does not automatically permit committing its responses to Git.
With an approved fixture, you can keep testing pure translation and reporting logic when the VPN is unavailable. This does not verify today's authentication, data freshness or pagination. Those need separate integration checks.
What changes when the nightly file becomes an API?
| Change | Likely location | Evidence needed before calling it compatible |
|---|---|---|
ORDER_NBR becomes orderNumber with identical identity rules | Source adapter | Contract and paired mapping tests |
| File becomes paginated API | Transport and integration checks | Complete enumeration, stable continuation behavior, limits and retry policy |
| Date becomes an instant | Domain contract and translation | Required business timezone and reporting cutoff |
closed now includes cancelled orders | Domain rules and reports | Agreed business meaning and revised expected outcomes |
| Data becomes available every minute | Workflow and operations as well as transport | Freshness measurement, load, access approval and ownership |
A seam contains representation changes. It cannot make a different business meaning equivalent, or turn an unapproved route into an approved one.
The spine above is what a record does as it crosses the seam, with the adapter marked as the one place allowed to know what the source meant.
Do this before moving on
Produce a one-page adapter contract and a passing script. Start with the table above, then add one status and one quantity edge case without weakening the existing checks. Change the fixture's status to 009: the current result must be rejection. If you decide it means closed, first write the new source-contract assumption, then add an explicit mapping and test.
Search an integration you know for raw status codes and field names. Separate expected occurrences in adapters, fixtures and tests from business calculations that depend on source representations. Describe one change you can isolate and one semantic change that must reach the domain model. A useful handover tells the next engineer both.
Go deeper
- Testability and dependency injection explains how to substitute a fixture-backed source without changing business calculations.
- Parsing messy data covers additional conversions and the evidence needed before normalizing values.
- Data quality connects rejected records to visible ownership and correction workflows.
- Enterprise ingestion across forty connectors extends the boundary problem to several source contracts and transports.
- The system nobody owns explains how to investigate source behavior when responsibility is fragmented.
Key takeaways
- Keep source representations behind an explicit translation contract.
- Preserve identifiers and business meaning; normalization needs evidence.
- Rejected values need an operational policy outside the pure conversion function.
- Offline fixtures require appropriate handling and cannot establish live compatibility.
- A transport upgrade may also change freshness, access, pagination or business semantics.
Check yourself
Answer before you look. Recalling it is what makes it stick; recognising it does not.
1The source contains customer IDs 00099421 and 99421. What justifies stripping zeros?
2The API replaces a date with a timezone-free timestamp. Is taking the first ten characters sufficient?
3Why can a nightly-file-to-API change require work beyond one adapter?
Sign in to track which lessons you have finished.
