Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Letters in the grader's text: A = MiniMax M3, B = Fable 5.1, C = muse, D = GPT-6.1 Sol, E = GPT-6 Luna, F = GPT-6 Astra, G = Opus 5.5, H = Grok 4.7, I = mimo, J = MiniMax M3.1 Flash, K = Sonnet 5.5.
B and G tie at 9.25. B is ranked first for a crisper purge-on-receipt design and a richer failure-path table; G is more concise and ships the lost-notes fix first. F and J tie at 8.75: F has stronger retention and risk reasoning, J better product judgment. D, H and C tie at 8.5. D's acceptance barrier is correct but heavy; H's protocol is the most implementable but rests on redefining 'accepted'; C is concise but leaves the offline purge gap as an exception path. K's length was weighed against it, but its added material is mostly substantive.
Evaluation 9.5 / 10graded blind as submission K
The most rigorous treatment of the hard problems. Purge has two layers: on learning of acceptance, and a local cap anchored on the server-stated submission-receipt time with monotonic-clock tracking. Because acceptance cannot precede receipt, the cap meets the deadline for a device that never comes back online. A purge interlock and a late-work path cover the rest. It ships durable drafts first and ends the pilot on scenario coverage rather than elapsed time. It is very long (about 11k words), and some scope (device hold, rescue export, residue tests) is heavy for a pilot.
Strengths
Section 10.2 proves the Layer-2 cap meets the deadline using server time, with handling for backward clock jumps and reboots
Design invariants I1–I9 are testable release blockers; scenario matrix of 20 cases
Milestone dependency graph with acceptance criteria per milestone; each decision has an owner and a 'needed by' milestone
Headline recommendations prioritize fixing data loss first
Weaknesses
Longest submission; some content repeats across sections
Pilot scope is heavy (device hold, rescue export, byte-level residue tests)
The late-work path deliberately deletes unsynced data at the deadline, which is reasonable but needs sign-off
Evidence the grader checked
§10.2 Layer 2: 'purging at submitted_at_server + R ... can never be later than accepted_at + 24h'
PLAN.md was written from BRIEF.md alone. I did not look at the existing app, APIs, or data, so every statement about how the current system behaves is an assumption. The plan lists sixteen (A1 to A16). These three carry the most risk, because each could invalidate a core design choice and none can be checked from the brief.
1. Concurrency premise: integer versions are reliable, and supervisor and inspector edits rarely overlap (A4 and A7)
What the plan assumes. The existing integer record version is a single counter per inspection, bumped by every writer (web, mobile, API, batch jobs). Stale writes are rejected atomically. Supervisors mostly edit a small set of fields, such as comments, corrections, and status, that seldom coincide with what an inspector is changing. On that basis, a field-level three-way merge plus an advisory device hold is enough, and no collaborative editor is needed.
Why it is risky.
If any writer path does not bump the version, a stale offline write can silently overwrite a supervisor's edit. That is undetected data loss, and it also corrupts the audit history.
If versions are coarser or finer than assumed, the merge design changes.
If overlap is common, conflict resolution stops being an edge case and becomes the product. That directly undercuts the brief's "no full collaborative editor".
Cheapest early test (M0).
Send a deliberately stale write through every writer path and confirm it is rejected.
Mine existing edit and audit history for supervisor edits made while an inspection was open elsewhere, to estimate the overlap rate.
If the assumption fails. Fix or wrap the offending writers before any offline client is enabled (M2). If overlap is high, tighten field ownership (D3) or add per-field locking before the pilot, and revisit whether the scope can stay non-collaborative.
2. Purge premise: "accepted" always follows a complete server copy, so the 24-hour rule can be met on offline devices without destroying work (A3 and A13)
What the plan assumes. Acceptance is a server event that can only follow a submission the server fully holds. That lets the tablet purge safely when it learns of acceptance. It also lets a local retention cap, counted from the server's stated receipt time, meet the deadline for a device that never hears about acceptance. The plan also assumes compliance accepts purging earlier than 24 hours after acceptance, and that server-side audit history is not "local customer data".
Why it is risky. Acceptance is learned by the tablet only when it has contact, and the tablets are the ones that lose contact. The design fails in two ways if the assumption is false.
If supervisors can accept or reassign an inspection while a tablet holds unsynced work, the regulation and the inspector's evidence collide. Complying means deleting work the server never received.
If compliance reads the rule more strictly, or requires retention of something we planned to purge, the whole purge design changes.
Related: the current read-through cache may already retain customer data past the limit (A16). That would be an existing compliance gap that must be escalated, not quietly fixed.
Cheapest early test (M0).
Read the accept code path and the state machine. Confirm whether acceptance requires a submission and whether it is reversible.
Get compliance's written answer on the deadline-passes-while-offline case (D1).
Inspect what the legacy cache holds today.
If the assumption fails. The plan has a late-work path for unsynced work, and it is the only place that permits deliberate data loss. That tradeoff needs explicit compliance and product sign-off (D1). If acceptance can occur without a full server copy, the server should warn or require a reason at acceptance when a device hold has unsynced work, or block it.
3. Evidence premise: offline-captured signatures and device-reported times are acceptable audit-grade evidence (A10 and A11)
What the plan assumes. A drawn signature captured offline, bound to a hash of the exact content signed, plus audit events that record both the untrusted device time and the trusted server receive time, satisfy the regulation's audit requirement and whatever legal weight the signature needs. It also assumes that a signed record can be frozen locally and that a later supervisor edit can be handled by marking the signature stale.
Why it is risky. The brief says a signature must be collected offline. It does not say whose signature it is, what it attests to, or what evidence regulators or courts require. If the requirement is a trusted timestamp, an identity check, or an on-device witness, the local flow cannot provide it. A tablet clock is user-changeable and cannot be trusted alone. Supervisors editing signed content after the fact is easy to under-specify. A wrong answer would only surface after the pilot has collected signatures that turn out not to be valid.
Cheapest early test (M0). Put three concrete questions to legal and compliance before designing the signature schema. Who signs, and what does it attest to? What time and identity evidence must be recorded? Is device-reported time plus server receipt time acceptable? Show them the exact audit record fields proposed in section 12.
If the assumption fails. Change the signature schema and possibly the flow: for example, require a server-confirmed time before a signature counts, or require a connectivity check at signing. This is why D6 and D11 are marked as needed before the signature milestone, and before schema work if they change stored fields. Nothing else in the plan depends on the answer, but the pilot must not collect real signatures until it is settled.
What I could not verify
Everything about the current stack: the mobile platform and framework (A12), the form-builder versioning (A6), whether the server supports idempotency or resumable uploads today (A15), and the actual photo sizes and tablet storage (A9). The plan puts each of these in M0 rather than guessing, and it states no dates, durations, or staffing numbers. Those depend on M0 findings and on team capacity that the brief does not provide.
sonnet-5.5/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
sonnet-5.5/06-product-planning/PLAN.md
Offline Inspection Sync: Product and Engineering Plan
Status: proposal for review. Input: BRIEF.md only. Nothing in the existing codebase or APIs was inspected, so every statement about current system behavior is an assumption (section 2) that Milestone M0 must confirm or overturn.
This plan deliberately contains no calendar dates, durations, or staffing numbers. Sequencing is expressed as dependencies and acceptance criteria. Durations and capacity can only be estimated after M0 answers the open questions and the team's capacity is known.
Contents
Summary and headline recommendations
Assumptions
User flow
Scope boundaries
Design invariants
Local data model
Sync protocol
Conflict policy
Attachments (photos and signature)
Retention and purge (the 24-hour rule)
Security and privacy
Audit history
Migration and backfill
Observability
Staged rollout
Rollback
Testing
Milestones
Decisions required before implementation
Brief-to-plan traceability
1. Summary and headline recommendations
Problem. Inspectors work in basements with no signal. Today, losing connectivity after opening an inspection can discard their notes. They need to open assigned inspections, fill forms, take up to 40 photos, and collect a signature with no connection. Connectivity comes back intermittently. Supervisors can edit the same inspection from the web. Regulation requires an audit history and removal of local customer data within 24 hours after acceptance. Product wants a pilot first and explicitly does not want a full collaborative editor.
Approach in one paragraph. Promote the tablet's SQLite database from a read-through cache to the system of record for any inspection the inspector has taken offline. Every edit is committed to SQLite before the UI acknowledges it and is appended to a durable outbox. A single-flight sync engine pushes outbox changes to the server as idempotent changesets guarded by the existing integer record version. Attachments upload separately, resumably, and verified by content hash. Conflicts are resolved per field with a three-way merge; the only conflict that reaches a human is the same field changed on both sides. Server-side audit is append-only. Local data lives in an encrypted store keyed so that purge is a deterministic, verifiable operation.
Headline recommendations (each is justified in its section and listed as a decision in section 19):
Fix data loss first. Ship durable local drafts (M1) before any offline sync. This addresses the brief's stated failure directly and de-risks everything after it.
Additive server changes only. Add idempotent changeset, submit-with-manifest, resumable attachment upload, and device-hold endpoints. Keep the integer version as the concurrency token. Do not change existing REST endpoints used by the web app.
Field-level three-way merge, not last-writer-wins, and not a collaborative editor. Supervisor-owned fields are server-authoritative. Same-field conflicts on inspector-owned fields are shown to the inspector, and no losing value is ever silently discarded.
Freeze on signature. A signature binds to a hash of the exact content signed. After signing, the local record is read-only. Changing signed content invalidates the signature.
Two-layer purge. Purge immediately when the device learns of acceptance, plus a local retention cap counted from the server's stated submission-receipt time. Acceptance can only happen after the server has the submission, so the cap meets the deadline even for a device that never hears about acceptance, as long as it knows the server received the submission.
Purge is guarded by a safety interlock. Local data is deleted only after the server has confirmed a matching manifest, except under an explicit compliance-approved override for the regulatory deadline (D1).
Pilot exit is coverage-based, not time-based. The pilot ends when specified hard scenarios have been observed and metrics pass, not after an arbitrary period.
2. Assumptions
The brief is short. These are the assumptions the plan rests on. "Risk" is the consequence if the assumption is wrong. "Validate in M0" is the cheapest way to find out. A1 to A16 are referenced throughout.
ID
Assumption
Risk if wrong
Validate in M0
A1
Inspections are created and assigned on the server. The tablet never creates one ("open assigned inspections").
Offline creation needs client-generated IDs and a create-conflict story; this plan does not cover it.
Product confirms.
A2
Tablets are MDM-enrolled, so we can enforce screen lock and disk encryption, disable OS backup, and remote-wipe. One inspector uses a tablet at a time under a personal login.
Shared or unmanaged devices break the security model (section 11) and per-user data separation.
Ask IT/MDM owner what is enforceable today.
A3
The server lifecycle is at least: in progress, submitted, accepted, with possible "returned for rework". "Accepted" is a discrete event stamped with a server timestamp, performed by a supervisor or system, and acceptance requires a prior submission.
The purge design (section 10) and the manifest interlock depend on it. If supervisors can accept without a submission, unsynced work can exist for an accepted inspection (D1).
Read the state machine and the accept code path. Check whether acceptance is reversible.
A4
The integer record version is one counter per inspection aggregate (form answers plus status), bumped by every writer (web, mobile, API, batch or import jobs), and stale writes are rejected atomically.
If some writer skips the bump, conflicts are silently missed and supervisor edits get overwritten. If versions are per-child or coarser, merge granularity changes.
Inventory every write path. Test with a deliberately stale write from each path.
A5
Some audit history exists today, but not necessarily field-level, device ID, device-reported time, or conflict-resolution provenance.
We may need to build most of the audit model (section 12), and it gates M2.
Sample existing audit records. Compare with what compliance requires.
A6
Form definitions are versioned and can change while a tablet is offline.
Answers captured against an old schema cannot be validated or stored server-side.
Check how the form builder versions and publishes changes.
A7
Supervisor web edits mostly touch a limited set of fields (review comments, corrections, status) and rarely the same fields an inspector is editing at the same moment.
If overlap is common, conflict UX becomes the product, and "no collaborative editor" is at risk.
Mine existing audit or edit history for supervisor edits made during inspector-open windows.
A8
Offline duration is unknown: from minutes (basement) to possibly multiple days (no coverage at site and no Wi-Fi afterward). The design must not depend on it.
Token lifetime, storage, retention, and lease policies all assume a bound.
Ask field ops for the realistic long tail. Instrument in M1.
A9
Tablet camera photos are multi-megabyte and tablet storage suffices for the configured number of simultaneous offline inspections. The 40-photo limit is per inspection.
Storage exhaustion during capture can corrupt work. Sizes drive upload feasibility.
Measure photo sizes and free storage on actual pilot tablets.
A10
The signature is a drawn signature captured on the tablet. Who signs (customer, site representative, or inspector) is not stated. It attests to the content at signing time. Legal requirements for e-signature evidence are unknown.
An unenforceable signature defeats the purpose of collecting it.
Legal/compliance answers D6.
A11
Device clocks can be wrong and users can change them. Audit needs a server-trusted time.
Device-reported times cannot stand alone as audit evidence.
Compliance answers D11.
A12
The app runs on managed iOS or Android tablets (platform not stated) with SQLite, and OS background-execution limits apply to sync and uploads.
Background upload and key-storage design differ by platform.
Confirm platform, framework, and minimum OS versions.
A13
"Local customer data" means everything on the tablet derived from an inspection: forms, photos, signature, thumbnails, caches, logs, exports. Server-side audit history is not "local customer data". The regulation sets a maximum retention, so purging earlier is permitted.
If compliance reads it more broadly or narrowly, the purge scope and the early-purge cap (D1) change.
Compliance confirms scope and that earlier purge is acceptable.
A14
Authentication uses expiring tokens with refresh. Offline token expiry policy is not defined.
Users could be locked out of their own unsynced data, or the reverse: stolen tablets stay usable indefinitely.
Read the auth design. Decide D9.
A15
The server has no resumable upload and no idempotency-key support today.
Retries during sync would duplicate photos or changes.
API inventory.
A16
The current read-through cache may already retain customer data past what regulation allows (it may never be purged).
An existing compliance gap surfaces as we touch the cache. It must be escalated, not silently fixed.
Inspect what the cache holds and for how long. Escalate to compliance if confirmed.
3. User flow
3.1 Inspector flow
Terminology used in the UI. These states must never be conflated, because "saved" must not be presented as "safe":
Saved on this tablet: committed to local SQLite.
Waiting to upload: local data that the server has not confirmed.
Uploaded: server has acknowledged the data and verified attachment hashes.
Submitted: server has the submission and its manifest matches.
Accepted: supervisor accepted. The tablet copy will be removed by a shown deadline.
Step 1. Prepare (online). While the tablet has a connection (office Wi-Fi, for example), the app pre-fetches assigned inspections in the prefetch window (D9). Each inspection shows a readiness badge: Ready offline means the record, form definition (pinned version, D10), reference data, and existing attachments needed to work are all stored locally. Not ready means something is missing. The inspector should see before entering the basement which inspections are ready. Opening a not-ready inspection with no connection shows an explicit reason and what is missing, never a spinner or a generic error.
Step 2. Work (offline or online, identical experience). Opening an inspection reads only from local storage. There is no "Save" button: every field change is committed to SQLite before the UI treats it as done. Connectivity changes never interrupt work, never show a blocking dialog, and never discard input. A persistent status chip shows the sync state ("Saved on this tablet", "Waiting to upload: 3 changes, 12 photos", "Uploaded").
Step 3. Photos. In-app camera only (never the system gallery, see section 9). The counter shows N of 40. The 41st capture is blocked with an explanation. Free storage is checked before capture starts and the inspector is warned when low.
Step 4. Signature. If the form requires one, the signature pad appears. On completion the app computes a hash of the form content plus attachment manifest as of that moment and stores the signature bound to it (D6). The inspection becomes read-only except an explicit "Edit and re-sign" action that discards the signature.
Step 5. Submit. Validation (required fields, required signature, photo minimums if configured) runs locally. "Submit" moves the inspection to submitted on this tablet. Sync sends it when a connection exists. The UI states plainly: "Submitted on this tablet. Not yet received by the office" until the server confirms, then "Submitted".
Step 6. Leaving the site. If anything is waiting to upload, the app tells the inspector how much and what will happen next ("2 inspections, 31 photos waiting. They will upload automatically when this tablet has a connection. Keep the app installed and do not clear its data.").
Step 7. Server accepts. When sync sees acceptance, the inspection shows "Accepted. Will be removed from this tablet by " and is purged (section 10).
3.2 Inspector edge flows
Situation
Behavior
App killed, tablet rebooted, or battery dies mid-edit
On next launch the draft is restored from SQLite. At most the field being typed at that instant can be lost, never earlier fields.
Supervisor changed different fields while the inspector was offline
Merged silently at next sync when the inspection is not on screen. A notice reads "Supervisor updated 2 fields."
Same field changed on both sides
Conflict card: both values, who changed each and when. Inspector chooses. Nothing is dropped (section 8).
Inspection reassigned or cancelled while offline
On contact, the device is told. If there are unsynced changes, they are sent as late offline work for supervisor review (D1). The inspection becomes read-only, then purges.
Supervisor returns inspection for rework
It reappears as editable, re-downloaded from the server at the new version, and the retention clock restarts.
Storage nearly full
Warning before capture. Below a hard floor, capture is refused with an explanation. The DB is never allowed to fail for lack of space.
Auth token expired while offline
Work continues under local unlock. Sync requires re-authentication when online. Local data is never wiped because auth failed.
Form definition changed on the server while offline
Inspector completes on the pinned version (D10).
Sync repeatedly fails for one inspection
The inspection is flagged "Needs attention" with a diagnostic code and a support action. It is never silently dropped or endlessly retried in a hidden loop.
3.3 Supervisor flow (web)
On an inspection with an offline copy, the supervisor sees a device hold banner: "Offline copy on 's tablet since
After sync, the supervisor sees field-level provenance: who set each value, whether it was device-reported offline, and any conflict resolutions.
Accepting an inspection while a device hold shows unsynced work needs an explicit confirmation with reason (D1). This is audited.
Purge status is visible: accepted, device purge confirmed, or overdue.
4. Scope boundaries
In scope for the pilot
Offline open of assigned inspections, form entry for existing form types, up to 40 photos, signature capture.
Durable local drafts (fixes the current data-loss failure).
Background sync with retry, idempotency, and integer-version concurrency.
Field-level merge and a minimal conflict UI.
Server-side audit extensions and purge receipts.
24-hour local data removal with verification.
Feature flags, observability, kill switches, and a staged pilot.
Minimal supervisor web changes: device hold banner, provenance display, purge status.
Out of scope (non-goals)
Real-time or collaborative editing, cursors, presence, CRDTs, operational transforms.
Offline creation of inspections (A1).
Offline supervisor or web use.
Video or audio attachments; more than 40 inspector photos; photo annotation or editing.
Device-to-device or peer sync.
Offline browsing of historical or completed inspections beyond the retention window.
Multiple tablets editing one inspection at once for the same inspector.
Unmanaged or personal devices.
Changing the REST style of existing APIs.
PDF or report generation offline (stays online unless M0 shows it is required).
Deferred candidates (revisit after pilot)
Thumbnail-first upload so supervisors see photos sooner.
Hard field-level locks.
Offline creation.
Bulk conflict tooling.
Time-boxed decryption keys for submitted data (extra defense for devices that stay offline and unpowered past the deadline).
5. Design invariants
These are testable rules. Violating any is a release blocker.
I1 Durability. No edit is acknowledged in the UI before its SQLite transaction commits (WAL mode, synchronous FULL for user-data writes).
I2 Outbox integrity. Outbox entries are append-only and are not deleted until the server has acknowledged them and the manifest verified.
I3 No silent loss. Any value that loses a merge, any op that fails validation, and any attachment that fails upload remains stored and visible to the user or support until resolved. Nothing is dropped without a user-visible or audit-visible record.
I4 Single deletion path. Customer data is removed only by the purge engine, which enforces preconditions and produces a receipt.
I5 Single root. All customer data reachable on the device belongs to exactly one inspection ID, so purge is "delete by root".
I6 Registered stores. Every table, directory, cache, and log that can hold customer data is registered with the purge engine. A CI check fails if a new store is added without registration.
I7 Idempotency. Every network mutation carries a client-generated ID. Replays produce the same result.
I8 Server authority. The server validates assignment, version, schema, and manifest on every write. The client is never trusted for authorization or time.
I9 Fail safe on flags. Turning a feature flag off never deletes unsynced data.
6. Local data model
The SQLite store changes from disposable cache to authoritative store for in-progress work. Legacy cache tables (prefix cache_) remain disposable. New tables (prefix local_) are never dropped or rebuilt without going through purge.
6.1 Tables
Table
Key columns
Purpose
local_inspection
id (server ID), server_version (last integer version merged into local state), lifecycle (see 6.2), assigned_to, form_schema_id, form_schema_version, downloaded_at, last_synced_at, submitted_at_local, submitted_at_server (server-stated receipt time, anchors the Layer 2 cap), accepted_at_server, purge_deadline, purge_target, readiness
One row per offline-capable inspection. Root of all customer data (I5).
local_base_snapshot
inspection_id, server_version, snapshot
Last server state the local state was merged against. Needed for three-way merge.
Durable record of purge work and receipts. Contains no customer content.
local_diag
ts, level, code, context_ids
Ring-buffer diagnostics, no customer content (allowlisted fields only).
Attachments are files in an app-private directory. They are never written to a shared gallery or external storage. Each inspection has its own directory and its own data-encryption key (D7).
Transitions are explicit, persisted, and each is idempotent so a crash mid-transition resumes safely.
6.3 Storage budgeting
Worst-case footprint per inspection is 40 x average_photo_bytes plus thumbnails and DB rows. With an illustrative 3 MB average that is about 120 MB per inspection, so ten simultaneously offline inspections would be about 1.2 GB. These are placeholders. M0 measures actual sizes and free space to set the maximum number of offline inspections per tablet and the free-space floor for capture.
6.4 Schema management
Forward-only, transactional, idempotent migrations with a version table.
Expand-then-contract: new tables and columns are added first and old ones are removed only in a later release. This is what makes app rollback safe (section 16).
Every migration is tested against fixture databases captured from each shipped schema version, including large and corrupted-but-openable ones.
7. Sync protocol
The protocol is illustrative and additive. Names are proposals for backend design review (D5). The existing integer record version is the concurrency token. Existing web-facing endpoints do not change.
7.1 Server surface (proposed)
Endpoint
Purpose
GET /sync/assignments?cursor=
Delta list: inspection ID, version, status, assignee, accepted_at, purge_deadline, revoked flag, form schema version. Returns server_time for clock anchoring.
GET /inspections/{id} (existing)
Full record and integer version for download.
POST /inspections/{id}/device-hold
Register or clear the advisory device hold on download, submit, or purge receipt.
POST /inspections/{id}/changesets
Idempotency key = changeset_id. Body: device_id, base_version, form_schema_version, ops with op_id, field_key, value, authored_at_device. Applies atomically against the version. Returns 200 {new_version} or 409 {current_version, current_record, conflicting_fields}.
POST /attachments/uploads and chunk, status, complete calls
Resumable, idempotent by client attachment ID. Server verifies SHA-256 on complete.
POST /inspections/{id}/submit
Body: final base_version plus a manifest (field-values hash, attachment IDs with SHA-256, signature hash). Server verifies everything is present, returns the missing list if not, and only then moves to submitted.
POST /devices/{id}/purge-receipts
Receipt of local removal (no customer content).
Server invariants:
Atomic version check plus write within one transaction. Each accepted changeset bumps the version by exactly one.
Idempotency store: a changeset with a previously seen ID returns the original response, not a second write.
Authorization on every call: the server re-checks that the caller is assigned to the inspection and the device is enrolled. Rejected changesets are quarantined and audited, not discarded.
Old clients and the web keep working. New behavior is opt-in per client capability.
7.2 Client sync cycle
Single-flight, interruptible, and every step persists its progress:
Preflight. Is a valid token present? Does a small authenticated request succeed? "Connected" from the OS is not trusted (captive portals, one-bar links). If preflight fails, back off with jitter and stop.
Pull deltas. Fetch assignment changes since the cursor and reconcile: new, updated, revoked, accepted. For inspections with local edits, store the new server snapshot as pending merge and never overwrite local state.
Push changesets. For each dirty inspection, oldest first, send changesets in seq order, one in flight per inspection. On 200, advance server_version and mark ops acked. On 409, go to the conflict path (section 8). On network failure, stop and retry later.
Upload attachments with bounded concurrency, independent of step 3 (section 9).
Submit any inspection in submitted_local once all its changesets are acked and its attachments verified, sending the manifest.
Purge queue (also runs on its own schedule with no connectivity, section 10).
Report purge receipts and diagnostics.
Persist cursor and schedule the next cycle.
Triggers: app foreground, connectivity-change events, a periodic background job, and a manual "Sync now". Sync never blocks editing.
7.3 Failure semantics
Response
Behavior
Timeout or network drop
Retry with capped exponential backoff and jitter. The idempotency key makes the retry safe even if the first request landed.
401 / token expired
Pause sync, prompt re-auth when the user is next in the app. Never wipe data.
403 / no longer assigned
Treat as revocation (D1 late-work path).
409 version conflict
Conflict path (section 8).
422 validation (for example schema mismatch)
Quarantine the changeset, mark the inspection "Needs attention", surface a diagnostic code. Do not retry blindly.
413 too large
Attachment path: shrink and retry once, then mark failed and visible.
5xx
Backoff. If persistent, stop after a bounded number of attempts per cycle and alert (section 14).
7.4 Clock handling
Device wall-clock time is recorded as "device-reported". Alongside it the device stores elapsed monotonic time since the last server contact, so the server can reconstruct approximate authoring time as last_server_time + elapsed. Ordering inside a device uses the per-inspection seq, never the wall clock. Backward clock jumps are detected and logged. The server stamps every audited event with its own receive time. Audit shows both (section 12).
7.5 Pull merges never disturb the open form
The engine will not mutate an inspection that is on screen. Remote non-overlapping changes are applied when the inspection is next opened or when the user accepts a "Supervisor updated this inspection" prompt.
8. Conflict policy
Premise. Product does not want a collaborative editor. The design therefore keeps concurrency rare by construction (field ownership, device hold) and handles the residue with a small, explicit policy. Frequency of overlap is unknown (A7); M0 estimates it.
8.1 Field ownership (D3)
Inspector-owned: observations, form answers, photos, signature. Web edits to these by supervisors are allowed but are corrections that can conflict.
Supervisor-owned: review comments, status transitions (return, accept), assignment. The mobile client renders these read-only and never writes them, so no conflicts are possible.
The server enforces ownership (I8), not just the UI.
8.2 Rules
Different fields changed on each side: automatic three-way merge using local_base_snapshot. Both authors recorded in audit.
Same field, same resulting value: no conflict.
Same field, different values (inspector-owned): a conflict is created. The server keeps its current value as current truth until the inspector decides. The inspector's device shows both values with actor and time. The inspector picks one (or edits a new value). The decision is sent as an ordinary changeset citing the conflict, and audit records the losing value.
Supervisor-owned fields: server wins, always.
Photos: append-only. A local delete is a tombstone (audit retains metadata; the blob is deleted from the device if it was never uploaded). A supervisor cannot silently remove an inspector photo without an audit event.
Status conflicts (accepted, cancelled, reassigned while the device holds unsynced work): server state wins. Unsynced device work goes down the late-work path (D1) and is never applied to an accepted record.
Submit racing a supervisor edit: the submit call returns 409. The inspector resolves and re-submits. A resubmit after resolution re-verifies the signature binding (section 9.4).
Retry bound: after a bounded number of consecutive 409s on one inspection, stop and mark it "Needs attention" rather than looping.
Nothing is silently discarded (I3).
8.3 Device hold (D4)
On download the server records an advisory hold: "copy on device X since T". Supervisors see it. It is informational and overridable because a lost or dead tablet must never block a record. It clears on submit, purge receipt, or expiry per policy. It reduces surprise, it does not prevent conflicts, and the merge policy does not rely on it.
8.4 Alternatives considered
Option
Why rejected
Last-writer-wins
Silently loses one party's work, and a device's "last" time is untrustworthy (A11). Fails the audit intent.
Whole-record hard lock
Blocks supervisors whenever a tablet is offline or lost. Used only as an advisory hold.
CRDT or operational transform
Explicitly a collaborative editor. Product declined it.
Supervisor always wins
Loses on-site facts only the inspector knows. Rejected for inspector-owned fields.
9. Attachments (photos and signature)
9.1 Capture pipeline
In-app camera only. Capturing through the system camera app would place photos in a shared gallery, outside the purge engine's reach and possibly in cloud backup. The app writes directly to its private, encrypted per-inspection directory.
Crash-safe write order: write to a temp file, flush, atomically rename, then insert the local_attachment row and the outbox op in one transaction. On startup, reconcile directory contents against rows: orphan files are quarantined and surfaced or cleaned up, and rows whose files are missing are flagged.
Processing: resize and compress at capture to a target set in D8 (starting parameters to be validated against inspection-quality requirements, not assumed). Strip or retain EXIF per D8. Store capture metadata in the database rather than depending on EXIF.
Limit of 40: enforced in the UI over non-removed inspector photos. The server enforces a separate, higher hard limit. If the combined total from all sources exceeds the cap at sync, the extra photos are stored and flagged, never dropped (I3).
Storage guard: check free space before entering capture. Refuse capture below a hard floor so the database always has headroom.
9.2 Upload
Metadata (attachment ID, hash, size, capture time) travels early in the changeset, so a supervisor can see "photo pending upload" before the bytes arrive.
Blobs upload through resumable sessions: create session, send chunks, query offset after any failure, resume, complete. The server verifies SHA-256 on complete, and only then is the local attachment marked verified.
Idempotent by client attachment ID plus content hash. Duplicate uploads collapse.
Bounded concurrency and chunk size tuned in M5 from measurements on real tablets and links. Network policy (Wi-Fi only versus also cellular on the company plan) is D8. Large uploads pause on low battery or low-quality links.
Per-attachment failure states are visible to the inspector and support, with retry. A failed attachment blocks submit and says which one.
9.3 Signature
Store vector strokes and a rendered image. Hash both. The image is uploaded as an attachment.
Record signer name and role (per D6), device-reported time, device ID, and the content hash.
9.4 Signature binding (D6)
At signing, compute a canonical hash over the form values, pinned schema version, and photo manifest. The signature record stores that hash.
After signing, the local inspection is read-only. "Edit and re-sign" invalidates the signature and returns to editing.
If a supervisor edit merges into signed content (or the server-side content hash no longer matches at submit), the signature is marked stale. The inspector must review and re-sign, or a supervisor decides. A stale signature is never presented as valid.
Whether the device-reported signing time is sufficient evidence is a legal question (A10, A11, D6, D11).
10. Retention and purge (the 24-hour rule)
This is the section where offline and regulation collide, so it is spelled out.
10.1 The problem
The regulation runs from acceptance, and acceptance is a server event. A tablet in a basement, or simply offline, may not learn of acceptance for a long time. A naive "delete on acceptance" therefore cannot meet the rule for offline devices. Conversely, deleting data that the server does not yet hold destroys the inspector's work and possibly evidence.
10.2 Design: two layers plus an interlock
Layer 1: server-driven. When sync reports accepted_at, the device sets purge_deadline = accepted_at + 24h and purges as soon as the interlock allows, which in practice is at that sync. Purging earlier than the deadline is permitted (A13). If the device learns of acceptance after the deadline, it purges immediately and records a late-purge event.
Layer 2: local retention cap (fallback for missed contact). When the server confirms a submission, its response (or any later pull) includes submitted_at_server, the server's own receipt time. The device stores it and schedules the local copy to be purged no later than submitted_at_server + R, where R is at most 24 hours (final value from compliance, D1). The anchor must be the server-stated receipt time, not the time the device happened to receive the acknowledgement. A lost response could otherwise delay the device's "synced" moment by hours, past acceptance. Acceptance cannot occur before the server received the submission, so accepted_at >= submitted_at_server, and purging at submitted_at_server + R with R at most 24h can never be later than accepted_at + 24h. The server-time deadline is converted to a local deadline using the clock anchor from the last contact (local_sync_state) plus elapsed monotonic time, tracked across reboots, and the stricter of wall-clock and monotonic elapsed time is used, with backward clock jumps detected and a safety margin applied. Layer 2 needs no connectivity once the device knows the submission was received. If the device learns of the receipt only after submitted_at_server + R has passed, it purges immediately and records a late-purge event. A device that does not know the server holds its submission is in the late-work situation (10.3), not in Layer 2. Returning an inspection for rework cancels the cap and restarts it at the next submission.
Interlock (safety). Before any purge, the purge engine verifies: all outbox ops acked, all attachments verified, and the server submission manifest matches the local manifest. If not satisfied:
If the inspection never reached submission, and the server has not accepted it, do not purge (the inspector's work is still needed).
If the inspection was accepted or revoked by the server while unsynced local work exists, follow the late-work path (10.3).
10.3 Late-work path (D1)
Cases: supervisor accepted or cancelled the inspection, or it was reassigned, while the device holds unsynced changes. Recommended behavior:
On contact, the device uploads the unsynced work as a late offline submission attached to the audit trail, not applied to the accepted record. A supervisor can review and, if needed, reopen or amend.
After the server confirms receipt, purge.
If the deadline passes before the device can make contact, compliance wins: the data is purged anyway. The event is recorded as "purged with unsynced data" (in the purge ledger and reported at next contact), the inspector is warned in advance ("This inspection was accepted. Unsent work will be removed at . Connect to send it."), and an alert fires. This tradeoff must be explicitly signed off by compliance and product (D1). It is the only place the plan permits deliberate data loss.
To make this rare, the server should warn or require a reason when accepting an inspection that has a device hold with unsynced work (section 3.3).
Files: photos, thumbnails, signature images, temp files, any exports.
Search indexes and derived caches.
Diagnostics and crash artifacts (which by rule contain no customer content, and are verified to).
Per-inspection encryption keys (destroyed).
Database residue: secure_delete, WAL checkpoint with truncate, and freed-page cleanup (vacuum) after purge.
OS-level surfaces controlled by MDM: backups disabled, screenshots and app-switcher previews blanked for customer screens, no customer data in notifications.
10.5 Verification and escalation
After purge, a verification pass confirms no rows, no files, no key.
Purge receipts (inspection ID, device, time, data classes, verified, had-unsynced-data flag) are queued and sent at next contact.
Server-side: any accepted inspection with no verified receipt after the deadline plus a grace period raises an alert and appears on a compliance dashboard. Escalation path: contact the inspector, then MDM remote wipe (D13).
The residue test (section 17) proves deletion at the byte level using planted marker strings.
10.6 Residual risks
A tablet that is powered off or unable to run the app past the deadline cannot purge. Mitigation: encryption at rest, device-bound keys, MDM lock or wipe. Deferred hardening: time-boxed decryption keys for submitted data.
A user who manipulates the clock. Mitigation: monotonic tracking and server-side detection at next contact.
11. Security and privacy
Threat
Controls
Lost or stolen tablet
Full-disk encryption (MDM-enforced). Database and attachment encryption with keys held in hardware-backed OS keystore (D7). App lock (PIN or biometric) with a timeout. Remote wipe via MDM. Server-side revocation blocks sync from that device.
Data left on device after acceptance
Purge engine (section 10). Registry of stores (I6) enforced by CI. Residue tests.
Data leaks via OS features
No gallery or external storage. OS backup disabled. Screenshot and app-switcher preview protection on customer screens. No customer content in notifications, logs, telemetry, or crash reports (schema-enforced allowlists).
Compromised or modified client
Server validates assignment, version, schema, ownership, manifests, and hashes. Device identity bound to an enrolled device. Rate limits.
Minimize what is prefetched: only assigned inspections within the prefetch window (D9), only the customer fields the form needs.
Offboarded or reassigned inspector
Assignment re-checked on each call. Revocation causes local read-only and purge on next contact. MDM wipe as escalation.
Stale offline credentials
Bounded offline session policy (D9). Sync requires fresh auth. Auth failure never deletes data (I9), but access after the bound requires re-auth.
Shared tablets
Assumed not to exist (A2). If they do, per-user data separation and logout behavior become a blocking decision. Logout with unsynced data must not destroy it: it locks the app and retains encrypted data.
Tampering with evidence
Content hashes over forms, photos, signature. Append-only audit (section 12).
Location or personal data in photo metadata
EXIF policy in D8.
Transport uses TLS. Whether to add certificate pinning is a platform and MDM decision recorded in M0.
12. Audit history
The regulation requires an audit history. Offline authoring complicates it: the time a change was made and the time the server learned of it differ, and the device clock is untrusted (A11).
Proposed server audit event (append-only, immutable):
event_type (field change, attachment added or removed, signature captured or invalidated, submit, accept, return, conflict opened or resolved, hold set or cleared, purge receipt, late offline submission, override)
field_key, old and new values (or hashes where content should not be duplicated)
version_before, version_after, changeset_id, conflict_id where relevant
Rules:
Audit is written in the same transaction as the change it describes.
Server-side audit is not "local customer data" (A13), but compliance must confirm both that it may retain customer content and for how long.
Deletion is audited: purge receipts and late-purge events go into the audit trail with no customer content.
Existing history is not fabricated. Records predating this feature are labelled as legacy with the fields they never captured (section 13).
The device's local outbox is the source of provenance and is purged with the rest.
13. Migration and backfill
13.1 On-device migration
Introduce the new schema by expansion only. Existing cache tables are untouched at first.
Cached rows are not converted into offline drafts. Their freshness is unknown. On first launch under the flag, assigned inspections are re-downloaded with versions.
Run the purge engine over legacy cache contents. Accepted or stale inspections in the legacy cache are customer data that may already be past the 24-hour limit (A16). If this is confirmed, escalate to compliance rather than fixing it silently.
Drop legacy cache tables only in a later release, after the new store has proven itself (expand-then-contract).
Migrations are transactional, idempotent, and reported to telemetry with outcome and duration. A failed migration leaves the old store readable and the app in online-only mode.
13.2 Server-side backfill and readiness
Item
Action
Integer versions
Verify every inspection has a non-null version and that every writer path bumps it (A4). Fix or wrap writers that do not, before any offline client is enabled. Backfill nulls with a baseline version.
Audit history
Add the new fields going forward. Mark pre-existing records legacy. Do not invent field-level history that was never recorded.
accepted_at
Ensure it exists for accepted inspections. Backfill from state history where available, otherwise leave null and never send a purge deadline for those (they cannot be on an offline device at launch).
Form schemas
Ensure schemas are versioned and old versions remain servable and acceptable for the pinned-schema window (D10).
Device registry
Create enrolled-device records mapped to MDM identities.
Idempotency store and upload sessions
New tables with expiry and cleanup.
Attachments
Nothing to backfill. Existing photos keep working.
13.3 Compatibility
Old app versions keep using existing endpoints. New endpoints are additive.
The server reports capabilities and minimum supported client version for offline features (D14). Old clients are never offered offline mode.
A new client against a server with offline features disabled falls back to online-only behavior while still draining any local data (I9).
14. Observability
Mobile telemetry must survive being offline: a local diagnostics ring buffer uploaded during sync. Telemetry uses allowlisted, schema-enforced fields and never contains customer content.
Signal
Why
Alert / use
Sync cycle success rate and failure codes
Health of the engine.
Alert on sustained drop.
Time from reconnect to fully synced
The user-visible promise.
Trend by cohort and site.
Outbox depth and age of oldest unsynced item per device
Detects stuck devices.
Alert when age exceeds a threshold set from pilot data.
Runbooks (M8) cover: stuck device, purge overdue, migration failure, suspected data loss, conflict spike, mass reauth, rollback.
15. Staged rollout
Stages advance on criteria, not on elapsed time. Every stage runs behind server-controlled feature flags (per org, user, and device). Flags are cached on device so they behave predictably offline, and fail safe (I9).
Stage
Who
Data
Entry
Exit
R0 Internal
Employees on test tablets
Synthetic customer data only
M3 to M8 functionality complete enough for internal use. R0 runs inside M9.
Full scenario matrix (section 17) passes on real devices, including real no-signal locations. Purge residue test passes.
R1 Closed pilot
Small cohort of volunteer inspectors and their supervisors. Cohort chosen so every manual sync can be reviewed. Includes basement-heavy and photo-heavy work and supervisors who edit often.
Real
R0 exit, compliance sign-off on D1, D6, D11, purge verified
Coverage-based (below) plus metrics green.
R2 Wider pilot
More sites and teams representing the variety in R1, plus known difficult cases
Real
R1 exit, fixes for R1 findings shipped
Same gates, plus support burden acceptable and no unresolved compliance findings.
R3 Staged expansion
Cohorts by org or site in increasing share
Real
R2 exit
Automated gates hold at each step (below).
R4 General availability
All eligible
Real
R3 complete
Online-only fallback kept for a defined period. Legacy cache code removed only after criteria are met.
Coverage-based pilot exit. R1 (and R2) are not complete until each of these has actually happened, either organically or as a scripted scenario if it did not occur naturally:
A real session with no signal for the whole inspection.
Flapping connectivity during upload.
A supervisor edit made while the inspector was offline, both non-overlapping and same-field.
A returned-for-rework cycle.
A full accept-to-purge cycle verified end to end.
A device reboot or app kill with unsynced data.
An app update with local data pending.
Low-storage handling.
A stale-signature case.
Metric gates (thresholds set in M0 from baseline, stated here as principles):
Zero confirmed silent data loss.
Purge compliance: every accepted pilot inspection purged within the deadline, or a documented, compliance-reviewed cause.
Sync completion after reconnect within the agreed target.
Conflicts resolved without support intervention above an agreed rate.
Inspector and supervisor feedback collected against a fixed script.
Automatic halt criteria. Pause expansion and page an owner on: any confirmed silent data loss; any purge overdue without explanation; migration failures above threshold; manifest-mismatch incidents; sustained sync-failure spike; audit gaps.
16. Rollback
Principle: rolling back must never strand or delete unsynced data.
Layer
Trigger
Action
Data safety
Irreversible
Feature flag
Serious defect in offline features
Drain mode: stop offering new offline sessions, keep syncing and purging existing local data. A separate hard "online-only" state applies only to inspections with no local data.
Unsynced data is preserved and continues to sync (I9).
Nothing.
App release
Defective build
Halt store rollout, ship a forward fix. Schema is expand-only, so an older build ignores new tables. An older build cannot sync local data, so the server enforces the minimum supported version (D14) and users with pending data are steered to update rather than downgrade.
Local data stays encrypted and intact.
Downgrading past the minimum version is unsupported.
Server
Sync endpoint defect
Disable the changeset path via flag. Clients queue and retry with backoff. Restore, then replay: idempotency makes replay safe. Additive schema means no data migration to undo.
Changes wait on tablets.
None, provided idempotency store retention outlasts the outage.
Purge logic
Purge deleting too much, or not enough
Interlock (section 10.2) is the first defense. Kill switch on the Layer 2 fallback cap only, with compliance on the loop. The server-driven purge (Layer 1) continues because it is a regulatory duty and is guarded by the interlock.
Deleted data cannot be recovered by design, which is why the interlock and the residue tests come first.
Deleted data.
Migration
Local migration bug
Old store remains readable, app falls back to online-only, forward-fix migration.
Legacy cache untouched.
None.
Stranded data
Device cannot sync (bug, lost credentials, app cannot open)
Support-initiated, admin-authorized rescue export of the encrypted local store, delivered over a controlled channel, for manual recovery (D13).
Preserves evidence.
Export is itself customer data and is purged per policy.
Rollback drills are part of M9: flag drain, forced update, sync endpoint outage with replay, and a purge-kill-switch exercise, all with test tablets holding unsynced synthetic data.
Property and model-based: a simulator with one server and multiple devices, random operations, random drops, duplicates, and reorderings. Property: all replicas converge; no accepted op is lost or applied twice; the version is monotonic; the audit log reconstructs every field's history.
Fault injection: kill the process or drop the network at every step of the sync cycle, the capture pipeline, the migration, and purge. Disk full. Token expiry mid-upload. Clock jumps forward and back. Reboot with unsynced data.
Contract tests between mobile and server: integer version semantics (409 shape), idempotent replay, manifest verification, upload resume, schema pinning. Run against every writer path (A4).
Migration: fixture databases from every shipped schema version; real-size data; interrupted migration; legacy cache purge.
Purge residue test: plant unique marker strings in fields, photos (in pixel-embedded and metadata forms), signature, logs, and caches; run the purge on a test device; scan the database file, WAL, free pages, and the app's whole data directory for the markers. No marker may be found. Repeat under crash-mid-purge.
Time-based purge: injectable clock to exercise both purge layers, the backward-jump case, and reboot elapsed-time tracking. One real-time soak for confidence.
Security: review of key handling and storage encryption, backup and screenshot behavior, MDM policy verification, authorization tests (wrong assignee, unenrolled device, replayed changesets), and independent penetration testing before R1.
End-to-end: web supervisor and tablet editing the same inspection in every interleaving from the conflict rules.
Field tests: real basements, real signal loss, flapping and captive-portal networks, low-battery and low-storage behavior on the actual pilot tablet models.
Performance and soak: 40-photo inspections over degraded links, several offline inspections queued, battery impact, long offline periods (A8), large local DB.
Usability: inspectors and supervisors run the scripted flows. Verify that the difference between "saved on tablet" and "uploaded" is understood.
Audit completeness: for a sampled set of test inspections, rebuild the full history from the audit trail alone and compare it with the known sequence of actions.
17.2 Scenario matrix (minimum set; each has an automated or scripted expected result)
#
Scenario
Expected
1
Connectivity lost right after opening an inspection, then 40 photos and a signature
All data present after reconnect. Zero loss.
2
App killed mid-typing
Everything before the current keystroke survives.
3
Sync request succeeds on server but response lost
Retry is idempotent. One version bump, no duplicates.
4
Supervisor edits a different field while offline
Auto-merge, both audited.
5
Supervisor edits the same field
Conflict shown. No value lost. Resolution audited.
6
Supervisor edits after signature
Signature marked stale. Cannot be submitted as valid.
7
Upload drops at 90 percent of a photo
Resumes from offset. Hash verified.
8
41st photo
Blocked with explanation.
9
Accepted while device online
Purged. Verified. Receipt received.
10
Accepted while device offline, later sync
Purged on next contact. Late-purge flagged if past deadline.
11
Submitted and synced, device never hears of acceptance. Variant: the submit response was lost and the device learned of receipt hours later
Layer 2 cap, anchored on submitted_at_server, purges within the deadline in both variants.
12
Accepted with unsynced device work
Late-work path. Data reaches audit if contact occurs. If the deadline passes first, purge with flagged event.
13
Reassigned or cancelled while offline
Read-only, late-work path, purge.
14
Form schema changed while offline
Completes on the pinned version.
15
Token expires offline
Work continues. Sync after re-auth. No wipe.
16
Clock set backward
Detected. Purge and audit unaffected.
17
Low storage during capture
Warned, then refused. DB intact.
18
Migration interrupted
Old store intact. Retry succeeds.
19
Feature flag turned off with unsynced data
Drain mode. Data still syncs.
20
Older app build with pending data
Server steers to update. Data preserved.
18. Milestones
Dependencies are hard prerequisites. Milestones that share no dependency may proceed in parallel. There are no dates. Each milestone ships its own metrics and tests; M8 assembles dashboards and runbooks from them.
M0 Discovery and decisions
| |
v v
M1 Durable local M2 Server foundations
store and purge (changesets, hold, audit,
primitives schema pinning, uploads)
| \ / |
| \ / |
v v v v
M3 Offline open M5 Attachments (needs M1 + M2)
and pull (M1+M2) |
| |
v |
M4 Push sync, merge, |
submit (M1+M2+M3) |
| \ |
| v v
| M6 Signature (M4 + M5)
| |
v v
M7 Retention and purge enforcement (M4 + M5 + M6)
|
v
M8 Observability and ops readiness (M4..M7 instrumentation)
|
v
M9 Hardening and pre-pilot verification
|
v
M10 Pilot (R1 to R2)
|
v
M11 Expansion and legacy retirement (R3, R4)
M0. Discovery and decisions
Depends on: nothing.
Work:
Answer every assumption A1 to A16 with evidence (code paths, API inventory, sampled data, device measurements).
Decide the blocking decisions in section 19 (see "needed by" column).
Write spike results: photo size and free-storage measurements on pilot tablets; supervisor-edit overlap estimate from history (A7); write-path and version audit (A4); audit-log capability gap (A5); MDM enforceability (A2); platform and background-execution constraints (A12); encryption library evaluation (D7).
Set metric thresholds and halt criteria (section 15) from baselines.
Acceptance:
Each of A1 to A16 is marked validated, invalidated, or unresolved with a named consequence and a plan change where needed.
D1, D2, D3, D5, D6, D7 are decided in writing by the named owners. The others have a decision or an explicitly agreed safe default with the milestone that needs the final answer.
A16 result reported to compliance if the legacy cache retains data past limits.
This plan is revised to reflect the findings.
M1. Durable local store and purge primitives
Depends on: M0 (D7, D3 for schema).
Work:
New local_* schema, migration framework, encrypted store, per-inspection key handling.
Every edit committed to SQLite before acknowledgement, and outbox append in the same transaction.
Draft restore on relaunch. Online behavior otherwise unchanged (this is the "stop losing notes" release).
Purge engine skeleton: store registry (I6) with CI check, delete-by-root, key destruction, verification, residue tooling.
Legacy cache purge on first launch (section 13.1).
Acceptance:
Fault-injection suite: killing the app or dropping the network at any point never loses a committed field (scenarios 1, 2, 18 pass).
Migration passes on all fixture databases, including interruption.
Residue test passes for the primitive purge.
CI fails when an unregistered store is added.
Shipped only behind a flag. No change in online functionality.
M2. Server foundations
Depends on: M0 (D2, D3, D5, D10, D11). Can proceed in parallel with M1.
Work:
Version audit fixes so every writer bumps the version (A4).
Idempotent changeset endpoint with atomic version check; idempotency store.
Minimal web changes: hold banner, provenance display, purge status, accept-with-unsynced warning.
Acceptance:
Contract tests pass for all endpoints, including duplicate delivery returning the identical result, stale version returning 409 with the documented body, and wrong assignee or unenrolled device being rejected and audited.
Every writer path passes the stale-write test.
Existing web and old-client regression suites pass unchanged.
Audit reconstruction test passes on sample flows.
Load test shows the idempotency store and upload sessions do not degrade existing endpoints.
M3. Offline open and pull
Depends on: M1, M2.
Work:
Prefetch of assigned inspections, readiness badges, local open with no connection, form definition pinning, revocation handling (read-only, then late-work or purge), status chip.
Acceptance:
With airplane mode enabled, a prefetched inspection opens and is fully editable; a non-prefetched one shows a specific explanation.
Reassign, cancel, and schema-change scenarios (13, 14) behave as specified.
Remote changes never mutate an inspection on screen.
Storage budget per tablet enforced from M0 measurements.
Submit state machine with manifest, "Needs attention" state, sync status UI.
Acceptance:
Scenarios 3, 4, 5, 15, 16, 19 pass.
Model-based simulator shows convergence and no lost or duplicated ops over many randomized runs including drops and duplicates.
Every 409 yields either an automatic merge or a stored conflict. Nothing is dropped.
The manifest at submit matches for all test inspections. The data-loss canary reports zero mismatches.
M5. Attachments
Depends on: M1, M2. Final integration with submit needs M4.
Work:
In-app camera pipeline with crash-safe writes, compression and EXIF policy, 40-photo counter, storage guard, encrypted per-inspection files.
Resumable upload client, bounded concurrency, network and battery policy, per-attachment failure states.
Acceptance:
Scenarios 7, 8, 17 pass. A 40-photo inspection completes on a degraded link in the fault-injection rig, including drops mid-chunk.
Files never appear in the device gallery or backup (verified on the pilot models).
Orphan-file and missing-file reconciliation tested after crash at each step of the write order.
Submit is blocked, with a specific message, while any attachment is unverified.
M6. Signature
Depends on: M4, M5 (D6 decided).
Work:
Signature capture, content-hash binding, read-only after signing, re-sign flow, stale-signature detection on merge and at submit, audit events.
Acceptance:
Scenario 6 passes. A signed inspection cannot be modified without invalidating the signature.
Hash computation is canonical and reproducible across client and server on the same content.
Legal or compliance reviewers approve the evidence recorded (D6, D11).
M7. Retention and purge enforcement
Depends on: M4, M5, M6 (D1 decided).
Work:
Both purge layers, interlock, late-work path, purge ledger, receipts, server alerts, compliance dashboard, user-facing deadline messaging.
Extend the registry to every data class introduced in M3 to M6.
Acceptance:
Scenarios 9 to 13 pass.
Residue test finds no markers across database, WAL, free pages, files, logs, and caches after purge, including a crash mid-purge.
Layer 2 cap demonstrably purges a submitted-and-synced inspection with no connectivity after acceptance, within the deadline, in time-injected tests and a real-time soak.
Backward clock jump and reboot tests pass.
Overdue-purge alert fires in a staged test.
Compliance signs off on the behavior in the deadline-passes-while-offline case.
M8. Observability and operational readiness
Depends on: instrumentation from M4 to M7.
Work:
Dashboards and alerts (section 14), diagnostics upload, support tooling, rescue-export procedure (D13), runbooks, on-call ownership.
Acceptance:
Each alert has been fired in a staging test and reaches a named owner.
Each runbook was exercised once by someone who did not write it.
Diagnostics contain no customer content (automated scan).
M9. Hardening and pre-pilot verification
Depends on: M3 to M8.
Work:
Full scenario matrix on real pilot tablet models, including physical basements and degraded networks.
Independent security review and penetration test, MDM policy verification.
Rollback drills (section 16).
Usability sessions with inspectors and supervisors.
Performance and soak.
Acceptance:
All 20 scenarios pass on real devices.
No open critical or high security findings.
Every rollback drill completes without unsynced-data loss.
Compliance sign-off recorded for D1, D6, D11 and the audit model.
R0 (internal, synthetic data) exit criteria met, which is what allows R1 to start.
M10. Pilot
Depends on: M9.
Work: run R1 then R2 (section 15) with regular review of the section 14 signals, the coverage checklist, and the halt criteria.
Acceptance:
Coverage-based exit list complete. Metric gates pass. No unresolved compliance findings.
Written findings: conflict rate versus A7, offline durations versus A8, photo and storage reality versus A9, support burden.
Decision memo on whether and how to expand, including any plan changes.
M11. Expansion and legacy retirement
Depends on: M10.
Work: R3 staged expansion with automated gates, R4 general availability, later removal of legacy cache code and the online-only fallback once criteria are met.
Acceptance:
Each expansion step holds the automated gates before the next.
Removal of legacy paths happens only after the agreed criteria hold and rollback is no longer required at that layer.
19. Decisions required before implementation
For each: the decision, options, a recommendation, who decides, and when it is needed. Owners are roles, not named people.
ID
Decision
Options
Recommendation
Decides
Needed by
D1
Purge rule, early-purge cap, and what happens to unsynced work at the deadline
(a) purge on learning acceptance only; (b) add a local cap counted from the server's submission-receipt time; (c) never purge unsynced data even past the deadline
(b). Layer 1 plus Layer 2 with interlock (section 10). For unsynced work on an accepted, cancelled, or reassigned inspection, use the late-work path; if the deadline passes before contact, purge anyway with a flagged event, because the regulation outranks retention. Compliance sets R (at most 24h). Reject (c): it violates the regulation.
Compliance/legal with Product
Before M1 schema is final; final by M7
D2
Conflict granularity and same-field policy
Record-level; field-level merge; last-writer-wins
Field-level three-way merge. Same-field inspector-owned conflicts go to the inspector, with both values kept. See section 8.
Partition: inspector-owned versus supervisor-owned, enforced on the server. Requires the form owner to classify fields.
Product
Before M1 schema
D4
Device hold
None; advisory hold; hard lease
Advisory, overridable hold, informational only. Hard locks strand records when tablets are lost.
Product
Before M2
D5
Sync API shape
Reuse existing full-record PUT; additive changeset, submit, and upload endpoints
Additive endpoints with idempotency keys and the integer version as concurrency token. Full-record PUT cannot carry provenance or idempotency.
Backend and Mobile leads
Before M2
D6
Signature semantics and post-signature edits
Store an image only; bind to content hash and freeze
Bind to a canonical content hash. Read-only after signing. Changes invalidate. Stale on merge. Legal confirms who signs and what evidence is required.
Legal/compliance with Product
Before M6 (before M1 if it changes schema)
D7
Local encryption and key management
OS file protection only; encrypted SQLite plus per-inspection file keys; per-inspection database files
Single encrypted SQLite (a maintained, platform-validated library) with a keystore-held key; attachments encrypted per inspection so purge can destroy the key; secure_delete, WAL truncate, and vacuum after purge. Per-inspection DB files are a viable alternative if M1 spikes show residue problems.
Security with Mobile lead
Before M1
D8
Photo policy: resize and compression targets, original retention, EXIF, network policy
Full-size originals; resized; Wi-Fi only versus cellular allowed
Resize and compress at capture to a target validated against inspection-quality needs. Do not retain separate originals unless required. Strip EXIF location unless required as evidence and record capture metadata in the database. Allow cellular on the company plan for small changesets and Wi-Fi-preferred for large uploads.
Product with Compliance
Before M5
D9
Offline scope and auth: what to prefetch, how long offline access lasts, when re-auth is needed
Prefetch all assigned; window only; on demand
Prefetch assigned inspections within a configured window. Bounded offline session length set from M0 data on offline duration. Sync always requires fresh auth. Auth failure never wipes data.
Product with Security
Before M3
D10
Form schema changes while offline
Force update; pin at download
Pin the schema version at download. The server accepts submissions on a previous version for a defined window. If incompatible, a supervisor migrates.
Product with Backend
Before M2
D11
Audit model and device-time trust
Trust device time; server time only; both
Record both with elapsed-since-contact. Server time is authoritative for ordering. Audit is append-only. Compliance confirms that this is acceptable audit evidence.
Compliance with Backend
Before M2
D12
Pilot definition: cohort, exit and halt criteria, metric thresholds
Time-boxed; coverage-based
Coverage-based exit (section 15). Cohort small enough for manual review. Thresholds from M0 baselines.
Product with Engineering and Compliance
Before M9
D13
Stranded-data recovery and remote wipe
No recovery; support-only rescue export; auto-upload of raw data
Admin-authorized, audited rescue export of the encrypted store. Remote wipe always available for security, but the server first shows "unsynced work at last contact" so recovery can be attempted.
Security with Support
Before M8
D14
Minimum supported version and old-client behavior
Allow all; capability negotiation with a version floor
Server-side capability negotiation. Old clients stay online-only and are never offered offline mode. Enforce a floor for users with pending local data.
Mobile and Backend leads
Before M2
20. Brief-to-plan traceability
Brief statement
Where addressed
Company-managed tablets, basements, no signal
A2, 3.1, 11
Lost connection can discard notes today
Section 1 (headline 1), I1, M1
Open assigned inspections, fill forms
3.1, M3
Up to 40 photos
Section 9, M5
Signature offline
9.3, 9.4, M6
Intermittent connectivity
7.2, 7.3, 9.2
Supervisors can edit the same inspection from the web
Section 8, 3.3, A7
Audit history
Section 12, D11
Local data removal within 24 hours after acceptance
Section 10, D1, M7
REST APIs with integer record versions
7.1, A4, D5
SQLite as read-through cache today
Section 6, section 13
Pilot before broad rollout
Section 15, M10
No full collaborative editor
Section 4 non-goals, 8.4
sonnet-5.5/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 9.25 / 10graded blind as submission B
A thorough and correct plan. It purges local customer data as soon as the server confirms complete receipt of the submission, so the 24-hour rule holds even for a device that stays offline, and every purge time comes from the server clock. The server does the merging by field class: the inspector wins on form content, the supervisor's value is kept and flagged, and there is no merge UI on the tablet. Milestones have hard dependencies, acceptance criteria and a stated critical path.
Strengths
D3 purge-on-confirmed-receipt meets the deadline by construction, with server-time purge_after, purge_confirmed receipts and an MDM backstop
Change-batch protocol with replay by batch_id, orphaned-work handling for cancel and reassign, and template pinning
Kill switch and rollback invariants that never discard queued work; client downgrade allowed only with an empty outbox
Every risky assumption is tagged 'verify in M0' with a stated consequence; the milestone table names a critical path and parallel tracks
Weaknesses
Long at about 7.5k words
Inspector-wins can override supervisor corrections; it relies on flags and the A7 log check
Purging at submission costs inspectors offline review (acknowledged)
Evidence the grader checked
§16 D3 recommendation and rationale
§9 Purge mechanism: 'set from server time in the response, never from the device clock'
RESPONSE — The three riskiest assumptions behind PLAN.md
PLAN.md labels every assumption (A#) and every pre-implementation decision
(D#). The three below carry the most risk because each one is load-bearing for
the whole design, each is unverified from the brief alone, and each would be
expensive to discover late. All three are scheduled for verification in M0
before any dependent milestone starts.
1. The integer record version is a real concurrency token (A3, feeds D1)
What is assumed. The version increments on every committed write to a
record, is returned on every read, and the server rejects writes with a stale
version, or can be changed to do so without breaking existing clients.
Why it is risky. The brief says only "REST with integer record versions".
That could mean an enforced optimistic-concurrency token, or an informational
counter that batch jobs bump, or a value that is not checked on write at all.
The entire merge design (§6.8, §7) uses the version as the base for deciding
which side changed a field. If the version is unreliable, the merge will
sometimes treat a supervisor's edit as unchanged and overwrite it without a
flag, or treat unchanged data as conflicting. That is silent data loss, which
is exactly the failure this programme exists to remove.
How to de-risk. In M0, read the server code paths that write inspections
(not the API docs) and produce a short report: where the version is incremented,
whether the check is enforced, and which write paths (web, batch, integrations)
bypass it. If enforcement is missing, it becomes the first task in M1 and the
plan already provides for a 409-impact report before it is switched on.
2. Supervisors rarely edit form content on in-progress inspections, so "inspector wins, supervisor is flagged" is acceptable (A7, D2)
What is assumed. Supervisor web edits during an inspection are mostly
workflow (reassign, reschedule, cancel, comment). Form answers are the
inspector's domain, so the policy applies the inspector's offline value and
records the supervisor's as overridden with a review flag.
Why it is risky. The brief rules out a collaborative editor, which forces a
policy-based merge. A policy is only acceptable if the losing side is rare and
the flag is actionable. If supervisors routinely correct answers on
in-progress inspections (for example fixing obvious errors during a quality
check while the inspector is still on site), the pilot will produce a stream of
review_required flags, supervisors will feel their work is being undone, and
the pressure will be to build the merge UI that product explicitly does not
want, or to flip the policy so the supervisor wins, which discards on-site
evidence. Either outcome changes the client, the server and the web scope.
How to de-risk. In M0, measure from server logs how often supervisor edits
touch form-answer fields on inspections in an in-progress state, and which
fields. If the rate is material, the cheap mitigation is on the web side: lock
inspector-owned fields while an inspection is in progress, or route corrections
through supervisor amendments (already designed for post-signature edits, D7).
Both keep the tablet free of merge dialogs. The pilot exit criteria (§12)
explicitly include supervisor judgement of flag volume.
3. Acceptance is a server-side event that can only follow complete receipt, and purging at submission confirmation is acceptable to the business (A5, A6, D3)
What is assumed. "Accepted" is a status transition made on the server,
with a server timestamp, and cannot happen before the server holds the full
submission including every attachment and the signature. The plan recommends
purging local customer data as soon as the server confirms complete receipt,
which is strictly earlier than acceptance and therefore always inside the
24-hour window.
Why it is risky. The regulatory clause is the one requirement in the brief
with an external penalty. A device in a basement, or powered off in a van, does
not learn that an inspection was accepted. Any design that waits for the
acceptance signal before purging can miss the 24-hour window through no fault
of the app. The recommended design avoids this entirely, but it depends on
acceptance never preceding complete receipt (otherwise the purge could happen
after the clock starts and while the device is offline), and it costs
inspectors the ability to review a submitted inspection offline. If product or
inspectors reject that trade, the fallback (D3 option b: purge on acceptance
signal with a timer from submission confirmation) reopens the offline-device
gap and turns a by-construction guarantee into an operational one that needs
server-side tracking and MDM backstops to be credible.
How to de-risk. In M0, confirm from server code whether acceptance can be
automated or applied to a partial submission, and get a written definition of
"customer data" from legal so the stub that remains on the device is agreed in
advance. Run the purge-timing scenarios from §14 (device offline for longer than
24 hours after submitting, device powered off, app force-stopped) in M5 under
both D3 options so the decision is made on evidence. Keep the server-side
"past purge_after without confirmation" alert regardless of the option
chosen; it is the only signal that reaches someone who can act.
Runners-up
Two further assumptions would hurt but are cheaper to recover from: A1
(assignments reach the device before signal is lost; if not, prefetch cannot
help and the product needs an out-of-band path) and A9 with D4 (compressed
photos are acceptable evidence; if not, storage and drain time on intermittent
links rise several-fold but the pipeline is unchanged).
fable-5.1/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
fable-5.1/06-product-planning/PLAN.md
PLAN — Offline inspections for the tablet app
Source of truth for product intent: BRIEF.md. Everything not stated there is an
assumption and is labelled A# in §2, with what breaks if it is wrong. Decisions
that must be made before implementation are labelled D# (§16). Milestones are
labelled M# (§15) and are ordered by dependency, not by calendar. This plan
deliberately contains no dates and no headcount figures; sequencing and exit
criteria are the commitments, elapsed time is not.
1. Goals and non-goals
Goals
ID
Goal
Derived from
G1
An inspector's work is never lost: not on signal loss, app kill, power loss, crash, reassignment, cancellation, or conflict.
"a lost connection ... can discard notes"
G2
An inspection can be completed end-to-end with no connectivity: open an assigned inspection, fill the form, take up to 40 photos, collect a signature, mark it complete.
brief, second sentence
G3
Sync happens opportunistically whenever connectivity appears, without the inspector doing anything; a manual "Sync now" exists but is never required.
"Connectivity may return intermittently"
G4
Supervisors keep editing from the web; concurrent edits are resolved by a fixed policy with a full record, not by a collaborative editor.
"Supervisors can edit the same inspection from the web", "does not want a full collaborative editor"
G5
Every change is auditable (who, what, when, from where, against which version), and local customer data is removed within 24 hours of the inspection being accepted.
"Regulations require ..."
G6
The feature can be piloted with a small group, observed, and switched off without losing queued work.
"Product wants a pilot before broad rollout"
Non-goals (see §4 for the full scope table): real-time co-editing, offline
creation of new inspections, offline reassignment, making the web app work
offline, video attachments.
2. Assumptions
Each assumption states why it matters and what changes if it is wrong. Items
marked verify in M0 are cheap to check and expensive to get wrong.
ID
Assumption
Why it matters
If wrong
A1
Inspectors have connectivity at some point between being assigned an inspection and arriving on site (office, vehicle, street level). Verify in M0.
The whole design is prefetch-then-work. "Open assigned inspections offline" is only possible if the assignment reached the device first.
Need an out-of-band delivery path (e.g. dispatcher phone call plus manual entry) or a paper fallback; that is a different product.
A2
An inspection has exactly one assigned inspector at a time; two tablets never edit the same inspection.
Conflict policy only has to handle inspector-vs-supervisor, not inspector-vs-inspector.
Same policy can be extended (device-level ownership), but the pilot must exclude multi-inspector jobs.
A3
The integer record version increments on every committed write to that record, is returned on every read, and the server rejects (409) a write that carries a stale version, or can be changed to do so. Verify in M0.
The version is the base for all merging. If it is advisory, or bumped by batch jobs, or shared across records, merge decisions will be wrong and silent data loss becomes possible.
Version enforcement becomes the first server task in M1 and nothing downstream can start until it is in place.
A4
Form templates are versioned and an inspection's template version is, or can be, pinned when the inspection is assigned.
The device must render the form it downloaded, and the server must interpret answers against the same template.
Pin at download time on the device and refuse to submit on mismatch; supervisor must re-issue the inspection.
A5
"Accepted" is a server-side status transition (by a supervisor or automation) that has a server timestamp and can only occur after the server holds the complete submission, including all attachments and the signature. Verify in M0.
Defines when the 24-hour purge clock starts and guarantees that purging never removes data the server does not yet have.
If acceptance can precede full receipt, purge timing must be tied to server receipt instead; see D3.
A6
"Customer data" means customer identity and contact details, premises address, premises photos, signature, and free-text fields that may contain personal data. A non-personal stub (inspection id, status, scheduled slot, template id) may remain on the device. Verify with legal in M0.
Determines what the purge removes and what the inspector still sees in the list afterwards.
If stricter: the stub holds only the id and status.
A7
Supervisor web edits to an in-progress inspection are mostly workflow (reassign, reschedule, cancel, comment) and rarely touch form answers. Verify in M0 from server logs.
The conflict policy (§7) favours the inspector for form content and flags the rest. If supervisors routinely correct form content, flags will be frequent and annoying.
Add a supervisor-side "review overridden edits" tool on the web (still no tablet merge UI); consider locking form content on the web while an inspector is in progress.
A8
Tablets are MDM-managed: device encryption, passcode, remote wipe, app-private storage, and background execution can be enforced by policy, and devices have free storage well above one day of inspections.
Security controls can lean on the platform; prefetch sizing can be generous.
Storage quotas become a hard cap on prefetch; every security control must be app-level.
A9
Photos are captured in-app; a size-limited, recompressed JPEG is acceptable evidence and originals need not be kept. Verify with legal in M0.
Sets storage and upload volume: 40 photos at roughly 1 MB each is a different problem from 40 photos at 5–8 MB each.
Keep originals; raise storage thresholds; expect longer drain times on intermittent links (D4).
A10
The server has, or can add, object storage with resumable multipart upload; failing that, the API tier can host a chunked upload endpoint.
Resumable upload is the only workable attachment path on intermittent links.
Build the chunked endpoint in M1; same client abstraction.
A11
Authentication is token-based with a refresh token whose lifetime can be set to cover the longest realistic offline period. Verify with security in M0.
Inspectors must open the app in a basement without contacting the identity provider.
Need a local unlock model with deferred re-authentication; must still never block draining queued work.
A12
The current app stores no user-authored drafts in SQLite (it is a read-through cache), so there is nothing to carry forward from old builds; the cache can simply be discarded. Verify in M0.
Migration is a drop-and-rebuild rather than a data migration.
Add a one-shot draft migration step to M2.
A13
The web app can be changed in the same programme (conflict flags, audit view, orphaned-work view).
Supervisors need somewhere to see and act on flagged merges.
Degrade to surfacing flags through existing comment fields and email notifications.
A14
Device clocks are not trustworthy. Server time is authoritative; device time and device-side sequence numbers are recorded for ordering.
Ordering of offline edits in the audit history.
Not an assumption that can be wrong, it is a stance; listed so nobody relies on device time.
A15
Regulators accept an audit history in which offline edits carry a device timestamp, a device-side monotonic sequence, and a server receipt timestamp. Verify with compliance in M0.
Defines what the change log must store.
Ask the regulator what ordering semantics they require before M1 finalises the event schema.
A16
Web edits go through the same server APIs and will therefore produce the same change events and version increments as tablet edits once M1 lands.
Merge and audit rely on a single change history.
Web write paths must be routed through the new change-event layer in M1; otherwise supervisor edits are invisible to the merge.
3. User flow
3.1 Prepare (needs connectivity)
Inspector opens the app on the network (office, vehicle). The app fetches the
assignment list and, for each assigned inspection within the prefetch window
(D14), downloads: the inspection record and its version, customer data, the
pinned form template, reference data the form needs (lookup lists), and
thumbnails of any existing attachments.
Each inspection in the list shows a readiness state: Ready offline /
Downloading / Not available offline (with a reason: storage, auth,
never fetched).
The inspector can mark additional inspections "Make available offline". The
app refuses only when storage is below the safety threshold and says so.
3.2 Work (no connectivity)
Inspector opens a Ready offline inspection. Every field edit is written
to the local store and the outbox in one transaction (autosave; text fields
debounced, flushed on blur, on navigation and on app background).
Photos are captured in-app, compressed, hashed, stored in app-private
encrypted storage, and linked to the form field they document. Counter shows
n / 40.
Inspector reviews the form, then Sign. Signature capture snapshots a hash
of the form content it attests. After signing, inspector-owned fields are
locked on the device; Edit again unlocks them and invalidates the
signature (must re-sign).
Mark complete transitions the local status to Complete, pending upload.
The inspector can move on to the next inspection.
3.3 Sync (whenever connectivity appears)
The sync engine wakes on connectivity change, app foreground, a periodic
background task, or "Sync now". It probes the API (not just the OS network
flag), then for each inspection with pending work: pushes change batches in
sequence order, applies the server's response (new version, server-side
changes, flags), then uploads attachments (resumable), then submits.
The list shows per-inspection sync state: n changes pending, uploading
photos 12 / 40, Submitted, Needs attention. Nothing about sync is
modal; the inspector keeps working.
3.4 After submission
Once the server confirms the complete submission (all attachments verified),
the device purges the inspection's customer data according to D3, leaving
a non-personal stub with status. When the supervisor accepts, the stub
shows Accepted. If the supervisor rejects or reopens, the inspection
reappears as an assignment and is re-downloaded on the next sync.
3.5 Failure paths
Situation
What the inspector sees
What the system does
App killed or tablet dies mid-edit
Reopens where they were, last debounce window at most lost
Transactional autosave; orphan-file sweep on start
Storage nearly full
Warning before capture; capture refused at hard floor
Prefetch paused; telemetry event
Supervisor edited a field the inspector also edited
Nothing on the tablet beyond a passive "Reviewed by supervisor" note
Inspector value applied, supervisor value kept in audit, inspection flagged on web (§7)
Supervisor cancelled or reassigned while offline
"This inspection is no longer assigned to you. Your work has been saved for review."
Changes stored server-side as orphaned work, not applied; local copy purged after confirmation
Upload keeps failing for one photo
"1 photo could not be uploaded" with retry
Retries with backoff; after threshold, surfaces to inspector and to observability; submission blocked until resolved or the photo is removed with an audited reason
Token expired while offline for longer than the policy window
"Sign in again when you have signal to open inspections." Already-open work remains viewable
Outbox drains automatically after re-auth; nothing is discarded
Template on server changed for a downloaded inspection
"Needs attention: this inspection was updated by the office"
Submission blocked; supervisor re-issues; work retained as orphaned if incompatible
4. Scope boundaries
In scope (this programme)
Explicitly out of scope
Possible later
Prefetch of assigned inspections; readiness indicator
Offline creation of new inspections
Offline creation with server-issued id blocks
Offline form editing with autosave
Editing inspections after acceptance (tablet or web)
—
Up to 40 photos per inspection, compressed, resumable upload
Video, audio, documents as attachments
Additional attachment kinds through the same pipeline
Signature capture bound to form content hash
Multiple signatures / countersign
Countersign as a second signature attachment
Queued, idempotent sync with field-level policy merge
Real-time co-editing, presence, CRDT/OT
Web-side "lock while in progress"
Supervisor web: conflict flags, audit history, orphaned work
Web app offline
—
Purge of local customer data; purge confirmation to server
Server-side data retention changes
—
Feature flags, kill switch, observability, pilot
Inspector-to-inspector handover offline
Device-level ownership in the same merge model
Force-upgrade floor after pilot
Support for non-managed (BYOD) devices
—
5. Local data model
Principles
The local store is authoritative for inspector-authored data until the server
acknowledges it. Server-derived data is a versioned mirror.
Every user write goes through the outbox in the same transaction as the data
write. There is no code path that writes user data without an outbox row.
Field-level answer rows rather than one JSON blob per inspection: this is what
makes field-level merge, cheap autosave, and per-field audit possible.
Client-generated UUIDs for anything created offline (change events,
attachments, batches, submissions). Server ids for inspections, which always
originate on the server.
No personal data outside the encrypted store (§9).
Schema (SQLite, encrypted per D5)
device_meta device_id, user_id, schema_version, last_purge_sweep_at
inspection id (server), server_version, template_id, template_version,
status_server, status_local -- see state list below
customer_json (PII), address_json (PII),
server_snapshot_json, -- last fully applied server state
downloaded_at, last_pull_at, last_push_at,
purge_after, purge_reason, review_flag_summary,
signature_attachment_id, signed_form_hash
form_answer inspection_id, field_id, value_json,
updated_seq, updated_at_device, -- last local change
server_value_json, server_version -- for display of "changed by office"
PK (inspection_id, field_id)
outbox seq (monotonic, device-wide), id (uuid), inspection_id,
entity ('answer'|'attachment_link'|'status'|'signature'),
op, field_id, old_value_json, new_value_json,
device_ts, actor_user_id, base_server_version,
batch_id (null until batched),
state ('pending'|'inflight'|'acked'|'orphaned'|'rejected'),
attempts, last_error, acked_at
attachment id (uuid, client), inspection_id, kind ('photo'|'signature'),
field_id, caption, local_path, sha256, bytes, mime,
captured_at_device, width, height,
upload_state ('pending'|'session_open'|'uploading'|
'uploaded'|'verified'|'failed'),
upload_session_json, -- session id, part size, done parts
server_attachment_id, attempts, last_error
submission id (uuid), inspection_id, form_hash, signature_attachment_id,
attachment_ids_json, created_at_device,
state ('pending'|'inflight'|'accepted_by_server'|'rejected'),
server_response_json
template id, version, json, fetched_at
reference_data key, version, json, fetched_at
sync_cursor scope ('assignments'), cursor, updated_at
audit_local mirror of acked outbox rows + server-originated events,
retained only until purge (for "history" view on device)
Inspection local states
not_downloaded → ready → in_progress → signed → complete_pending_upload
→ submitted → purged_stub (then accepted / rejected reflected on the
stub). Side states: needs_attention (with reason), orphaned (work saved on
server, no longer assigned).
Retention on device
outbox rows in state acked are kept until the inspection is purged, then
removed; they double as the local audit mirror.
Attachments in verified state are kept until purge (so the inspector can
review them) unless storage pressure forces earlier eviction of verified files
only (never unverified ones).
6. Sync protocol
6.1 Principles
Ship intent, not state. The device sends change events (field, old, new,
base version), never a whole record. The server merges.
Everything is idempotent. Batches, upload sessions, submissions and purge
confirmations carry client-generated ids; the server stores results and
replays them on retry.
Small before large. Form changes (kilobytes) are pushed before
attachments (megabytes) so that a ten-second window of signal still gets the
inspector's notes to the server.
Never block editing on the network. Sync is a background concern; the UI
only reflects its state.
One ordering per inspection. Per inspection: push, then apply the
response, then pull if needed, then upload, then submit. Different
inspections may proceed in parallel; the same inspection never has two
in-flight operations.
Server time is authoritative. Device timestamps are recorded, not trusted.
6.2 API surface (all additive; existing endpoints unchanged)
Endpoint
Purpose
Idempotency
GET /me/assignments
Assignment list with inspection_id, server_version, status, template_version. Cheap; polled.
—
GET /inspections/{id}?include=customer,template,attachments
Full download for offline use; returns version.
—
POST /inspections/{id}/change-batches
Push a batch of change events with base_version. Server merges per §7, records events, bumps version once per applied batch, returns outcome per event plus server-side changes since base_version.
batch_id; server stores responses for replay
POST /inspections/{id}/attachments
Open an upload session for a client-generated attachment_id (sha256, bytes, mime, kind). Returns the upload mechanism (presigned multipart parts, or chunked endpoint) and part size.
attachment_id
GET /inspections/{id}/attachments/{aid}/upload-status
Which parts the server has. Used on resume; server is truth, not the device's own record.
—
POST /inspections/{id}/attachments/{aid}/complete
Server assembles, verifies sha256, marks verified.
attachment_id
POST /inspections/{id}/submissions
Submit with submission_id, version, form_hash, signature_attachment_id, attachment_ids. Server checks all attachments verified, template match, status submittable.
submission_id
POST /devices/{device_id}/events
Device telemetry and compliance events, including purge_confirmed.
Existing PUT /inspections/{id} with a version check remains for the web and
old app builds; in M1 it is rerouted internally through the same change-event
layer so that web edits produce events (A16).
Outcomes: applied, applied_with_flags, not_applied (with reason:
not_assigned, cancelled, template_mismatch, status_locked; the batch is
stored as orphaned work), rejected (malformed; never for concurrency).
6.4 Push algorithm (per inspection)
Select pending outbox rows for the inspection in seq order up to a batch
size cap; assign batch_id; mark inflight (transaction).
base_version = inspection.server_version.
POST. On network failure: revert rows to pending, back off. On 5xx: same.
On 200: mark rows acked (or orphaned), set inspection.server_version = new_version, apply server_changes_since_base to server_snapshot_json
and to form_answer for fields with no pending local rows, store flags.
On app restart with rows in inflight: re-send the same batch_id; the
server replays the stored response. This is why the batch id is persisted
before the first attempt.
If more pending rows exist (the inspector kept editing), loop.
6.5 Pull algorithm
GET /me/assignments on every sync wake. For each assignment:
not local → download (prefetch window and storage permitting);
local and server_version newer and no pending outbox rows → GET and
apply;
local with pending rows → do nothing; the push response will carry server
changes.
Local inspections absent from the list → orphaned or unassigned; push any
pending rows once (server will store as orphaned work), then schedule purge.
Full-record apply overwrites server_snapshot_json, form_answer for fields
without pending local rows, status, template version, and the customer data.
6.6 Submission
After all outbox rows are acked and all attachments verified, POST the
submission. 409 with missing_attachments → re-verify attachments (the
device's record may lag the server); 409 version cannot happen if push ran
first, but is handled by pushing again; 409 not_submittable → needs_attention.
6.7 Triggers, backoff, connectivity
Wake on: OS connectivity change, app foreground, periodic background task,
capture or edit while online, manual "Sync now".
Reachability = a HEAD /health with a short timeout, not the OS "connected"
flag (captive portals, one-bar signal). Failed probe = treat as offline,
retry on next trigger.
Backoff: exponential with jitter per inspection, reset on any success.
Attempts and last error are stored so support can see them.
Attachments: one at a time per device, parts sequential, progress persisted
per part; on resume ask the server which parts it has.
6.8 Versions with an integer record version
The server increments version by exactly one per applied batch (not per
event) and records every event with the resulting_version.
The audit log is the same store the merge reads: "server changes since base
version N" is a query over events with version > N. This is deliberate: the
compliance requirement and the merge requirement are satisfied by one table.
Records first touched after M1 get a baseline event at their current
version so history is reconstructible from that point (see §10).
7. Conflict policy
Design constraint from the brief: no collaborative editor. Therefore: a fixed,
documented, field-class ownership policy, applied by the server, with every
losing value retained in the audit history and surfaced to the supervisor on
the web. The tablet never shows a merge dialog.
7.1 Field classes
Class
Examples
Who edits on device
Rule when both sides changed since base
Inspector-owned
form answers, photos and captions, inspector notes, signature
inspector
Inspector value applied; supervisor value recorded as overridden; review_required flag on the inspection
Supervisor-owned
assignee, scheduled slot, cancel/accept/reject, supervisor comments, customer record corrections
nobody (read-only on tablet)
Server value applied; device overwrites local copy on next sync
System
version, status derived from transitions, timestamps
nobody
Server
Every field in every template is mapped to a class in a table owned by product
(D2). A field without a mapping is inspector-owned if it is on the form and
supervisor-owned otherwise; the mapping test in CI fails if a template field is
unmapped.
7.2 Merge procedure (server, per event)
Given an event with base_version = B for field F:
If no server-side event for F with version > B exists → apply.
Else if F is inspector-owned → apply the inspector's value; record the
supervisor's value as overridden; add review_required.
Else (supervisor-owned) → do not apply; record the event as superseded;
the response tells the device the server value.
Structural transitions are checked before field events:
Server state when batch arrives
Result
Assigned to this inspector, not locked
Merge as above
Cancelled, or reassigned to someone else
not_applied; batch stored as orphaned work attached to the inspection's audit; supervisor sees it and can reattach it when re-issuing
not_applied, status_locked, alert raised in observability
7.3 Signature interplay
The signature attaches to a form_hash of inspector-owned content at signing
time. The server stores that hash.
If a supervisor edits an inspector-owned field after signature (web), the
server records the edit as a supervisor amendment with its own audit
entry and marks signature_covers = false for that field. Supervisor
amendments do not rewrite signed content silently. Whether amendments require
re-signing by the inspector is D7.
If the inspector unlocks and edits after signing, the device invalidates the
local signature and requires a new one; the old signature remains in the
audit history.
7.4 What each party sees
Inspector: a passive note when supervisor-owned fields changed (e.g. the slot
moved) and when their work was orphaned. No decisions to make offline.
Supervisor (web): review_required badge with a list of overridden values,
the audit history, and orphaned work sets with a "reattach to re-issued
inspection" action.
8. Attachments
Capture pipeline (single transaction where it touches the database):
capture to a temp file → recompress to the configured maximum long edge and
quality (D4) → compute sha256 → move to the app-private attachment directory
(excluded from OS backup) → insert attachment row + outboxattachment_link event. On startup, sweep the directory for files with no row
(crash between move and insert) and delete them.
Size envelope: 40 photos per inspection. At the recommended compression
target this is on the order of tens of megabytes per inspection; without
compression it can be several hundred. Prefetch and storage thresholds (D14)
are set from the chosen target.
Upload: open a session, then upload parts; persist each completed part;
on resume ask the server for its part list; complete → server verifies the
hash → verified. Only verified attachments count toward submission. Upload
order is capture order; the signature is last because it is captured last.
Storage management: a pre-capture free-space check with a warning threshold
and a hard floor; prefetch pauses at the warning threshold; verified attachment
files may be evicted under pressure (re-downloadable), unverified ones never.
Signature: stored as an attachment of kind signature (image plus stroke
data) with the form_hash it attests; same upload path.
Removal: an inspector may delete a not-yet-submitted photo; this is an
audited attachment_link remove event and the file is deleted locally; if the
attachment already reached the server, the server keeps it in the audit store
and marks it removed from the inspection.
9. Security and compliance
At rest on device
SQLite encrypted (SQLCipher or equivalent) with a key held in the OS keystore
(D5). Attachment files under OS file-based encryption with the strongest
data-protection class compatible with background upload; app-level per-file
encryption if security requires (D5).
Database and attachment directories excluded from OS cloud backups and from
device-to-device transfer. This is a test in M5, not a checklist item.
No customer data in logs, crash reports, telemetry, notifications, or the
recent-apps screenshot (secure-flag / snapshot masking).
Camera writes to app-private storage only; never to the shared gallery.
Authentication offline
Refresh token in the keystore with a lifetime covering the offline window
policy (D6). App open offline is gated by device unlock plus an app-level
biometric/PIN check bound to the keystore.
After the policy window with no successful refresh, the app requires online
re-authentication to open inspections but keeps the outbox and drains it
after re-auth. Nothing is discarded on auth expiry.
Authorisation on server
Every batch, upload session and submission is re-checked against assignment
at application time, not at download time. Work on inspections no longer
assigned is stored as orphaned, never applied, never dropped.
Batch replay by id is tenant- and user-scoped.
Purge of local customer data (D3 sets the trigger; this is the mechanism)
purge_after is set per inspection when the server confirms it holds the
complete submission (or the orphaned work, or an unassignment). It is set from
server time in the response, never from the device clock.
Purge runs: immediately on receiving the confirming response (device is
online at that instant), on every app foreground, on the periodic background
task, and before rendering anything on cold start. It deletes customer and
address JSON, form answers, attachment files, acked outbox rows and local
audit mirror; it leaves the stub. Files are deleted, then the SQLite
secure_delete pragma is relied on for row overwrite.
The device posts purge_confirmed {inspection_id, purged_at_device, purge_after} to the server. The server tracks, per device, inspections past
purge_after without confirmation and raises an alert; the backstop is an
MDM action on that device. A powered-off device is a documented residual
risk, minimised by choosing the earliest purge trigger (D3).
Audit history
Every change event stores: actor, device id, device timestamp, device
sequence, server receipt timestamp, base version, resulting version, old and
new values, merge outcome, and for overridden values the competing value and
actor. Server-side event store is append-only. Web edits produce the same
events (A16).
Purge events, orphaned-work events, signature events and submission events
are part of the same history.
Transport: TLS; certificate pinning is D5's second half (recommended for a
managed fleet, low cost).
10. Migration and backfill
Client
Schema version bump. On first launch after upgrade: verify A12 (no drafts in
the old cache), drop the cache tables, create the new schema, force a full
prefetch on next connectivity. A device that upgrades while offline shows
"Connect to download your inspections" — this is the only moment the new app
needs signal that the old one did not; the rollout instructs inspectors to
upgrade on the network (MDM can enforce Wi-Fi-only install).
Downgrade protection: the app refuses to run against a newer schema; MDM
must not push an older build to a device with pending outbox rows (§13).
Server
New tables: change_event (append-only), change_batch (id → stored
response), attachment_upload_session, submission, orphaned_work,
device, device_event, inspection_flag. All additive; no column drops in
this programme.
Reroute the existing PUT /inspections/{id} and web write paths through the
change-event layer so every write produces events and a version increment.
Backfill: no historical events are synthesised. Each inspection receives
a baseline event (current version, current snapshot) the first time the
new layer touches it. Audit history is complete from that point; earlier
history remains wherever it lives today. Compliance sign-off on this boundary
is part of M0.
Version enforcement: if A3 turns out to be advisory today, add the check and
run a report of how many current write paths would now receive 409s before
enabling it.
Web
Show review_required flags, overridden values, orphaned work, audit
history; no other behaviour change.
Coexistence
Old app builds keep working against the unchanged endpoints throughout the
pilot. A minimum-version floor is enforced by remote config only after the
expand stage (§12) has met its exit criteria.
11. Observability
Client events (no personal data; ids and counts only), buffered locally and
shipped with sync:
Sync: attempts, outcomes by error class, batch size, round-trip time,
outbox depth and age of oldest pending row per device.
Attachments: bytes uploaded, parts resumed, sessions abandoned, verify
failures, time from capture to verified.
Merge: flags by field class, orphaned batches, template mismatches.
Purge: time from purge_after to purge_confirmed; purge failures.
Readiness: time to "Ready offline", prefetch failures by reason, storage
headroom, captures refused for storage.
A ring-buffer diagnostic log (no personal data) that support can request from
a device through remote config.
Server metrics: batch apply latency, outcomes distribution, replayed batch
ids (a proxy for flaky links), upload session completion rate, submissions
rejected by reason, devices past purge_after without confirmation, orphaned
work volume, 409 rate on the legacy PUT after version enforcement.
Dashboards and alerts for the pilot
Per-device "oldest pending change" above a threshold.
Any inspection past purge_after + margin without purge_confirmed.
Upload failure rate or verify-failure rate above baseline.
Stages are gated by exit criteria, not by elapsed time. Controls: server-side
per-user offline_enabled flag, client remote-config kill switch, minimum app
version.
Stage
Who
Entry
Exit criteria
S0 Dogfood
Engineering and product on a test tenant with network shaping
M7 complete
Fault-injection suite green on real devices; purge SLA met in every scripted scenario; no unsynced-work loss in a full day of shaped-network use
S1 Pilot
One team of inspectors and their supervisors, real inspections
S0 exit; supervisors briefed on flags; support runbook in place
Over the pilot window: zero loss-of-work incidents; every completed inspection eventually submitted and accepted through the offline path; purge confirmed within SLA for every accepted inspection; flag volume and supervisor feedback judged acceptable by product (A7 check); no open severity-1 defects
S2 Expand
Additional teams, including at least one with a different connectivity profile
S1 exit
Same criteria at larger scale; support load sustainable without engineering on call for the feature
S3 General availability
All inspectors
S2 exit
Minimum-version floor enforced; legacy cache code path removed in a subsequent release
The kill switch behaviour is fixed: turning offline mode off stops prefetch
and new offline edits but always keeps draining the outbox and uploading
attachments. There is no configuration that discards queued work.
13. Rollback
Invariants that hold under every rollback: no queued work is discarded; no
purge is skipped; server-side data stays consistent with the audit log.
Layer
Rollback action
Constraint
Feature flag / kill switch
Disable per user or globally
Outbox continues to drain; the app reverts to online-only for new work
Client build
MDM rolls back to the previous build
Only permitted for a device whose outbox is empty and attachments are all verified; the app exposes this via a device event so the MDM operator can check. Otherwise the previous build cannot read the new schema and would strand work
Server endpoints
Flag off the new endpoints
Devices see 404/410 and back off; nothing is lost; must not be combined with a client rollback that leaves work stranded
Server schema
Not rolled back; all changes are additive and unused tables are harmless
Column or table drops are deferred until after S3
Merged data
Supervisor reverts an overridden value from the audit view
Every merge stored the losing value, so revert is a normal edit with its own event
14. Testing
Unit — merge engine (property-based: single-side change always applies;
both-side change resolves by class; replaying a batch is a no-op; event order
within a batch is respected), outbox state machine, purge scheduler with an
injected clock, storage threshold logic, form-hash stability across
serialisation.
Contract — client and server against a shared API contract for every
endpoint in §6.2, including the replay behaviour of every idempotent call.
Integration with fault injection — a harness that cuts the network at
every state transition: mid-batch, between batch and response, mid-part, after
complete but before the response, during submit; plus 5xx bursts, slow links,
token expiry mid-sync, server clock ahead/behind device. Pass condition: every
run ends with server state equal to device intent and no duplicated or lost
events.
Device — app kill during capture; power loss (pull the battery on a test
unit) during autosave; airplane mode toggled every few seconds; background
execution restricted by the OS; storage filled to the floor; OS upgrade with
pending work; restore-from-backup must produce no customer data.
Concurrency scenarios — scripted supervisor edits against each field class
while a device is offline with competing edits: verify applied value, audit
entries, flags, and what each party sees. Include cancel, reassign, template
change, and accept-before-push (must be impossible; test that it is).
Attachment soak — 40 photos across multiple inspections on a shaped
intermittent link; verify resumption counts, integrity, and total drain time.
Security and compliance — at-rest encryption verified on a jailbroken/rooted
test unit; purge verified by forensic file-system inspection; log and telemetry
scan for personal data; audit reconstruction test: rebuild any inspection's
state at any version from events alone.
Pilot acceptance script — a checklist inspectors and supervisors run in the
first days of S1 covering the full flow and the failure paths in §3.5.
15. Milestones
Dependencies are hard unless marked (soft). Acceptance criteria are the
definition of done; nothing else is.
ID
Milestone
Depends on
Deliverables
Acceptance criteria
M0
Discovery and decisions
—
Verified answers for A1, A3, A5, A6, A7, A9, A11, A12, A15; decisions D1–D15 recorded; field-class mapping for every current template; compliance sign-off on the audit baseline boundary (§10)
Every D# has a recorded owner and choice; no "verify" assumption remains unverified; server version semantics documented from code, not from memory
M1
Server foundations
M0
Version enforcement; change_event layer with web/legacy writes rerouted; change-batch endpoint with merge and replay; flags; attachment sessions; submissions; device events incl. purge confirmation; remote config; audit endpoint
Contract tests pass; audit reconstruction test passes; legacy PUT 409 report reviewed; merge property tests pass for every field class
Fault-injection suite green; server state equals device intent in every run; no duplicate events under replay
M4
Attachments and signature
M2, M1 (sessions)
Capture pipeline; compression; resumable upload; verify; storage management; signature with form hash; removal
40-photo soak passes on a shaped link with resumption; integrity verified server-side for every file; storage floor respected
M5
Purge and compliance
M3, M4
Purge scheduler and triggers; purge confirmation; server-side purge tracking and alert; audit completeness on device and server
Purge SLA met in every scripted offline scenario; forensic check finds no residual customer data; alert fires for an unconfirmed purge in test
M6
Supervisor web
M1
Flags and overridden values; audit history view; orphaned work with reattach
Supervisors in dogfood can find, understand and act on every flag type without engineering help
M7
Pilot readiness
M3, M4, M5, M6
Dashboards and alerts (§11); kill switch verified; support runbook; inspector and supervisor briefing; pilot acceptance script
S0 exit criteria met
M8
Rollout S1 → S3
M7
Pilot operation, expand, GA, version floor, legacy path removal
Stage exit criteria in §12
M1 and M2 can proceed in parallel against the agreed contract. M4 and M3 can
proceed in parallel once M2's schema is stable. M6 is independent of the
client track. The longest chain is M0 → M1 → M3 → M5 → M7 → M8.
16. Decisions required before implementation
Each has a recommendation; the recommendation is what this plan assumes.
ID
Decision
Options
Recommendation and rationale
D1
Where merging happens
(a) server-side field-level merge via a change-batch endpoint; (b) client-side three-way merge over the existing whole-record PUT
(a). Whole-record PUT with a version check turns every supervisor touch into a 409 for the whole inspection; the server already needs the event log for audit, and the merge is a query over it. (b) is a fallback only if the server cannot change in this programme.
D2
Conflict ownership policy
(a) inspector wins on form content, supervisor wins on workflow, losers audited and flagged; (b) supervisor always wins; (c) last write wins by server time
(a). The inspector was on site; (b) discards field evidence silently; (c) is arbitrary under intermittent connectivity. Product owns the field-class table.
D3
When local customer data is purged
(a) immediately when the server confirms the complete submission (or orphaned work / unassignment), keeping a non-personal stub; (b) on the acceptance signal with a fallback timer from submission confirmation; (c) only on the acceptance signal
(a). Compliance holds by construction: purge happens while the device is online and before acceptance can occur, so the 24-hour window is never at risk from a device that stays offline. The cost is that inspectors cannot review a submitted inspection offline; a rejected inspection is re-downloaded. Choose (b) only if the pilot shows inspectors need post-submission access, and set the fallback timer well inside 24 hours.
D4
Photo handling
(a) recompress on capture to a fixed maximum edge and quality, keep only the compressed file; (b) keep originals and upload them
(a), subject to legal confirming compressed images are acceptable evidence (A9). Storage and drain time on intermittent links scale directly with this choice.
D5
At-rest protection and transport
(a) SQLCipher for the database plus OS file encryption for attachments, backup exclusion, certificate pinning; (b) add app-level per-file encryption for attachments
(a) for the pilot on a managed fleet; escalate to (b) if security review requires it. Backup exclusion is non-negotiable in either case.
D6
Offline authentication window
Refresh-token lifetime and local unlock policy
Set the window to exceed the longest realistic offline period found in M0, gate app open on device unlock plus app biometric/PIN, and require online re-auth after the window. Draining the outbox never depends on the window.
D7
Post-signature edits by supervisors
(a) amendments recorded separately with signature_covers=false, no re-sign required; (b) any amendment to signed fields invalidates the signature and returns the inspection to the inspector
(a) for the pilot; it never rewrites signed content and keeps the flow simple. (b) is a compliance question for legal; if required it is a status transition on the server, not a client change.
D8
Offline creation of inspections
in / out
Out for this programme. It needs server-issued id ranges and a different assignment model.
D9
Upload mechanism
(a) presigned multipart to object storage; (b) chunked upload through the API
(a) if object storage exists, otherwise (b); the client sees one session abstraction either way (A10).
D10
Pull mechanism
(a) poll the assignment list with versions; (b) server change feed
(a). Assignment lists are small; polling on each sync wake is cheap and simple. Revisit if the list or the fleet grows.
D11
Work on cancelled or reassigned inspections
(a) store as orphaned work, never apply, supervisor can reattach; (b) apply anyway; (c) discard
(a). (b) corrupts a record the inspector no longer owns; (c) violates G1.
D12
Version floor and legacy removal
when to enforce a minimum app version
After S2 exit; never during the pilot, so that a rollback to the old build remains possible for devices with an empty outbox.
D13
Template pinning
(a) pin template version at assignment, server never changes it for a downloaded inspection; (b) allow template changes and migrate answers
(a). Template changes on an in-progress inspection are re-issues, not edits.
D14
Prefetch window and storage thresholds
how many days of assignments; warning and hard-floor free-space values
Prefetch today's and the next working day's assignments automatically, allow manual "Make available offline" beyond that, and derive the thresholds from the D4 compression target and the device fleet's free storage measured in M0.
D15
Ordering semantics for the audit history
device sequence + device time + server receipt time (A15)
Record all three; present server receipt order by default with device order available. Confirm with compliance in M0.
17. Open questions (owner to be assigned in M0)
Does acceptance ever happen automatically, and can it happen on a partial
submission? (A5)
What exactly counts as customer data for the purge, and may the stub retain
the scheduled slot and a non-identifying title? (A6)
How often do supervisors edit form answers on in-progress inspections today?
(A7; answerable from server logs)
Is a recompressed image acceptable evidence? (A9)
What is the longest period a tablet realistically stays offline? (D6, D14)
Is the audit baseline boundary (history complete from M1 onward) acceptable
to the regulator? (§10)
Are there inspection types that legitimately have more than one inspector?
(A2; excluded from the pilot if so)
18. Risk summary
The three assumptions whose failure would most change this plan are A3 (version
semantics), A7 with D2 (conflict policy acceptability), and A5 with D3 (purge
timing under offline conditions). They are discussed in RESPONSE.md. Each is
scheduled for verification in M0 before any dependent work starts.
fable-5.1/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 9.25 / 10graded blind as submission G
Strong and practical. Purge happens on server-confirmed receipt with a backstop timer that never deletes unsynced data. The server does a per-field merge using field categories. Work on reassigned or cancelled inspections is quarantined rather than dropped. An integrity canary detects lost ops. Durable local drafts ship first as M1, which fixes the lost-notes problem early, and a dependency graph shows parallel tracks.
Strengths
D3 purge-on-receipt reasoning, plus a server alert for accepted inspections without a purge receipt
Server-side field merge with base_value, op dedup, and receipt_hash verification
M1 ships the lost-notes fix on its own; the milestone graph marks M1, M2 and M5 as parallel
Rollback includes an audit replay tool and an emergency path to export unsynced data
Weaknesses
The backstop purge deadline is left as 'maximum retention since last server contact' for compliance to set; the collision with unsynced data is only escalated
Inspector-wins may override supervisor corrections, mitigated only by flags
PLAN.md covers the full plan for offline inspection work: assumptions A1–A12,
scope, decisions D1–D11 with recommendations, user flow, SQLite model, sync
protocol, conflict policy, attachments, security, migration, observability,
rollout, rollback, testing and milestones M0–M9. The plan rests most heavily on
the three assumptions below. Each one must be validated in M0, before any build
work is committed.
1. The 24-hour removal rule can be met by purging on server receipt (A1, D3)
Assumption. Deleting local customer data as soon as the server confirms full
receipt of a submission satisfies the rule, with a local timer as a backstop. The
rule does not require the device to prove deletion within 24 hours of acceptance.
Why it is risky. Acceptance happens on the server, often while the tablet is
offline. A device that goes back underground right after submitting cannot learn
about the acceptance in time. The design only works if purging before
acceptance is allowed. It also depends on how "local customer data" is defined:
logs, OS backups and stubs holding only an ID might count. And it depends on what
happens when the rule conflicts with the other requirement of never discarding
unsynced field work. If the interpretation is wrong, this is a regulatory breach
on every such device, not just a bug.
Validate by getting a written interpretation from compliance/legal in M0, and
testing a device that goes offline immediately after submitting in M5. A purge
receipt report on the server makes compliance measurable.
2. Forms can be merged field by field, and same-field conflicts are rare (A2, A8, D2)
Assumption. Form templates have stable field IDs, and repeating sections can
be given stable row IDs. Supervisors seldom edit the same field an inspector is
editing offline. Given that, a per-category rule is enough ("inspector wins on
fields captured on site, supervisor wins on review and status fields, the losing
value is kept and flagged").
Why it is risky. The whole conflict policy, and the promise of "no full
collaborative editor", depend on it. If records are stored as positional arrays
or free-text blobs, field-level merge is not safe without a data migration. If
concurrent same-field edits are common, a flag on the web will not be enough.
Supervisors would need a real review queue, and the pilot scope would grow.
Validate by auditing the template and record schema, and measuring from
production edit history how often the web and mobile clients write to the same
record, and which fields.
3. The server can add an idempotent field-level changes API without disrupting existing clients (A3, D1)
Assumption. A new batch /changes endpoint can be added next to the existing
REST API with integer versions. The endpoint handles op-ID deduplication, merge on
the server and audit writes. The web editor and other clients keep working
unchanged, and all of them write through a shared audit path.
Why it is risky. Connectivity is intermittent, so "committed but the response
was lost" is the normal case, not an edge case. A record-level PUT with an
integer version cannot tell a lost acknowledgement apart from a real conflict. It
also cannot record device time or operation IDs for the audit history the
regulation requires. If the server cannot change (ownership, capacity, or the web
editor bypassing the version or audit), the fallback is rebasing on the client.
That leaves weaker duplicate protection and an incomplete audit trail, which
threatens both the "no lost work" goal and compliance.
Validate by a server codebase and ownership review in M0, and by confirming
that every existing write path increments the version and can emit audit events.
opus-5.5/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
opus-5.5/06-product-planning/PLAN.md
Plan: Offline Inspection Work and Sync
Status: draft for review. Source: BRIEF.md (the only product input). Anything
not stated in the brief is labelled as an assumption (A#) or an open decision
(D#). The plan contains no calendar dates or staffing numbers. Sizing should be
done by the implementing team after Milestone 0, once the decisions in section 3
are settled.
0. Problem statement and goals
Problem. Inspectors work in basements with no signal. If the connection drops
after an inspection is opened, notes can be lost. The mobile app uses SQLite only
as a read-through cache, so work in progress is never stored durably on the
device.
Goals
No lost work. Anything an inspector enters is saved on the device before
the UI confirms it, and survives connection loss, app kill, reboot and low
battery shutdown.
Full offline task. Inspectors can open an assigned inspection, fill in the
form, take up to 40 photos, collect a signature and submit, all without a
connection.
Correct sync over intermittent connectivity. Sync resumes on its own,
applies each change exactly once, and handles concurrent supervisor edits from
the web by a documented policy. It never loses data silently.
Compliance. Every change has a complete audit history. Local customer data
is removed within 24 hours after an inspection is accepted.
Safe delivery. A pilot comes before broad rollout, with a kill switch that
never strands unsynced work.
Non-goals (from the brief: "does not want a full collaborative editor"):
real-time co-editing, character-level merge of text, live cursors or presence,
and multi-user editing on the device.
1. Assumptions
Each assumption lists how it will be checked and what changes if it turns out
false. Assumptions marked high risk are covered in RESPONSE.md.
#
Assumption
How to validate (M0 unless noted)
If false
A1
High risk. The 24-hour removal rule is satisfied if local customer data is deleted once the server has confirmed full receipt of the submission, which happens before acceptance. It is also satisfied if a device that stays offline deletes the data on a local timer. The regulation does not require the device itself to prove deletion within 24 hours of acceptance.
Written interpretation from compliance/legal; confirm what counts as "customer data" and "local" (for example whether app-private encrypted storage and OS backups count).
If the device must prove deletion, add MDM-driven remote wipe and a compliance report of unconfirmed devices. The rule then also limits how long a device may stay offline.
A2
High risk. Inspection forms have stable field identifiers. Repeating sections (defects, rooms) can be given stable row IDs, so merging field by field is meaningful.
Inventory the form template schema and the server record shape; check how the web editor saves (whole record or partial).
If forms are free-text blobs or positional arrays, the model must change (row UUIDs, template migration) before sync can merge safely. That makes M2 larger.
A3
High risk. The server team can add an idempotent batch "changes" endpoint and audit writes alongside the existing REST API with integer versions. Web and other clients keep using the existing endpoints unchanged.
Review the server codebase and ownership; confirm that the web editor increments the same integer version on every write.
If the server cannot change, fall back to client-side rebase on the existing PUT + version (D1, option B). Idempotency and audit are weaker in that case, see D1.
A4
Tablets are company-managed through MDM, which can enforce OS encryption and a screen lock and can push app updates and remote wipes.
Confirm the MDM product and policies with IT.
Add in-app encryption with an app PIN and our own wipe-on-command handling.
A5
Inspectors know which inspections they will do before they lose signal, because the inspections are assigned in advance.
Interview pilot inspectors; look at assignment lead times in production data.
Add bulk "download today's route" and a background prefetch of assignments. Creating inspections offline stays out of scope.
A6
A typical offline period lasts hours, not days, with signal between sites (vehicle, street level).
Pilot telemetry (longest offline period, age of oldest unsynced change); inspector interviews.
Raise the offline auth grace period (D4) and storage budget; review the purge timer (A1).
A7
Tablets have enough free storage for at least a day's inspections with 40 photos each after on-device compression.
Inventory device models and free storage via MDM; measure compressed photo sizes.
Stricter compression, a per-inspection storage check before starting, and upload of photos early when signal appears.
A8
Supervisors rarely edit the same field an inspector is editing at the same time. Most concurrent supervisor edits are to review, status or comment fields.
Analyse server edit history: how often the web and mobile clients write the same record within an inspection window, and which fields.
If same-field conflicts are common, a conflict review queue becomes pilot scope, not a follow-up.
A9
Form templates change rarely, and an in-progress inspection may stay on the template version it was downloaded with.
Confirm the template versioning process with product/ops.
Add server-side template migration for submitted records.
A10
The customer signature is a signature on the inspection as captured at signing time, not on later supervisor edits.
Compliance/legal and product.
Signature must be re-collected after material edits, which is not feasible offline and after the visit. This needs a product decision (D7).
A11
Device clocks can be wrong. Server receive time is authoritative for ordering. Device time is recorded as evidence only.
Given. Verify skew in pilot telemetry.
None. The design already does not trust device time.
A12
The existing mobile SQLite cache holds nothing that cannot be fetched again. No user-authored data lives only in the cache today (the brief says notes are discarded, which suggests they are held in memory).
Code review of the current app's persistence.
The migration must carry any drafts it finds across before dropping cache tables (see section 9).
2. Scope
In scope (pilot and GA)
Download assigned inspections for offline use, with their form template version
and any reference data the form needs (pick lists, prior findings if shown).
Filling in forms offline, including repeating sections.
Up to 40 photos per inspection, taken in the app, with captions and links to
fields.
Customer signature capture offline.
Submitting offline: the submission is queued and completes when all data and
attachments reach the server.
Background sync that resumes when connectivity returns.
A field-level conflict policy with supervisor web edits (section 6).
An audit history of every change, both device-originated and web-originated.
Local purge of customer data within the required window.
A sync status UI: pending changes, pending photos, last successful sync, and
errors that need attention.
A feature flag, kill switch and telemetry.
Out of scope (candidates for later phases)
A full collaborative editor, real-time co-editing, and merging inside a single
text field.
Creating new, unassigned inspections offline.
Reassigning, cancelling or scheduling inspections on the device.
Offline mode for the supervisor web app.
Video, or audio notes.
Photos imported from the device gallery. Camera capture in the app only, so
customer photos never land in shared storage.
Changing the form template in the middle of an inspection.
A general conflict resolution UI on the tablet. Conflicts are shown to
supervisors on the web. Inspectors see only a notice.
3. Decisions required before implementation
Each decision has options and a recommendation. Decisions D1 to D5 and D7 block
M2 or M5 and must be settled in M0.
#
Decision
Options
Recommendation
Blocks
D1
How writes reach the server
A: New idempotent batch endpoint POST /inspections/{id}/changes that takes field-level operations with base values and does the merge on the server. B: Client-side rebase on the existing PUT with version (on 409, fetch, merge locally, retry).
A. Option B cannot tell "my earlier request was applied but the response was lost" apart from a real conflict. It also cannot record device time, operation IDs or conflict decisions in the audit trail, and every client would need merge logic. A is additive, so existing REST clients are unaffected.
M2, M3
D2
Same-field conflict rule
A: Last write at the server wins. B: Supervisor always wins. C: Rule set by field category, with the losing value kept and flagged.
C. Fields the inspector captures on site (observations, measurements, photos, signature) are decided by the inspector's value, because the inspector was on site. Supervisor-owned fields (status, review notes, assignment, due dates) are decided by the supervisor's value and cannot be edited on the device. The losing value is always kept in the audit record and raises a conflict flag on the web. Option A lets whoever syncs last win, which in practice is always the inspector coming out of a basement. That is arbitrary.
M3
D3
When the device purges customer data
A: When the device learns the inspection was accepted. B: Right after the server confirms full receipt of the submission (record plus all attachments), with a local timer as a backstop.
B, subject to A1. Under A, a device that stays offline after submitting cannot learn about acceptance, so it could break the 24-hour rule. Under B the device purges before acceptance could happen. If a supervisor returns the inspection for rework, the device downloads it again.
M5
D4
Offline authentication
A: Require a valid access token for every action. B: Allow offline work for a limited grace period after the last online login, with a device-bound refresh credential and an OS or app unlock.
B, with a grace period set by security (suggested as a question: "one working shift plus margin"). Unsynced work is never deleted when the grace period ends. The app locks and requires online re-authentication before continuing.
M5
D5
Encryption at rest
A: Rely on OS full-disk encryption managed by MDM. B: Also encrypt the database (SQLCipher or equivalent) and photo files with keys held in the Keystore/Keychain.
B. It protects data if the device is unlocked but the app sandbox is extracted. Crypto-erase (destroying the key) is also a quick backstop for purge. Measure the performance cost on the lowest-spec pilot tablet during M1.
M1, M5
D6
Photo size and format
Full-resolution originals, or compression on the device to a maximum dimension and quality.
Compress on the device to a resolution agreed with supervisors and compliance, as the smallest size that is still usable as evidence. Keep the EXIF capture time and strip GPS unless it is required. Record a SHA-256 hash after compression.
M4
D7
What the signature attests
A: The whole inspection as later accepted. B: A content snapshot at the moment of signing.
B (see A10). Compute a canonical hash of the inspector-captured content at signing and store it with the signature. Later edits by the supervisor are audited separately. If an inspector edits after signing, the app asks whether to collect the signature again.
M4
D8
Server-side state changed while offline (reassigned, cancelled, template retired, already accepted)
Reject the device's changes, or accept them into quarantine.
Accept into quarantine: never throw away field work. The server stores the incoming changes against the record in a quarantined state and flags the record to supervisors. The device is told the outcome, and purge goes ahead once the server has the data.
M3
D9
How inspections are made available offline
Manual "make available offline" per inspection, or automatic download of all assigned inspections for the next N days.
Automatic download of assigned inspections for a limited window, plus a manual option. Show the storage impact. This avoids relying on inspectors remembering (A5).
M3
D10
Signalling to supervisors that an inspection is being worked offline
No signal, an advisory "checked out on device" badge, or a hard lock.
Advisory badge showing the time of the last device sync. No hard lock, because a lock held by an offline device cannot be released and this is not a collaborative editor.
M3
D11
Pilot cohort and exit criteria
Choose a region or team, and success thresholds.
One team that regularly works below ground, on the lowest-spec device model in use. Exit criteria are in section 12.
M7, M8
4. User flow
Inspector (tablet)
Before going offline (at the office, in the vehicle or on street-level
signal): the app syncs in the background. Assigned inspections within the
download window (D9) are downloaded together with their template and
reference data. Each inspection card shows its offline readiness: ✓ ready,
downloading, or not available offline.
Open an inspection without a connection. It opens from the local database.
If an inspection was not downloaded, the app says so clearly instead of
showing a spinner.
Fill in the form. Every field change is written to SQLite in a transaction
before the UI acknowledges it. The app autosaves continuously and has no
"save" button.
Take photos, up to 40. Each photo is compressed, encrypted, hashed and
saved before the next shot is allowed. The counter shows "n / 40". The app
checks free storage when an inspection starts and warns early.
Collect the signature. The customer signs. The app stores the signature
image and vector strokes together with the content hash (D7).
Submit. Validation runs locally against the template rules. If validation
passes, the inspection moves to submitted_pending_sync and becomes read-only
on the device. The inspector can move on to the next site.
Connectivity returns, even briefly: the outbox drains in order and photos
upload in resumable chunks. The status bar shows "Syncing 12 changes, 18
photos", then "All changes saved to server".
Server confirms full receipt. The device receives a receipt for the
operations and all attachment hashes, then purges the customer data (D3) and
keeps only a non-identifying stub (inspection ID and "submitted at").
Conflicts or quarantine. The inspector sees a non-blocking notice such as
"A supervisor also edited 2 fields; they will review". The inspector has no
decisions to make on the device.
Supervisor (web)
Sees an advisory "checked out on device, last synced …" badge (D10).
Can edit as today. Edits to supervisor-owned fields always apply. Edits to
fields the inspector captures apply, but may be superseded when the inspector's
offline change syncs. The superseded value stays visible in the conflict panel.
Reviews flagged conflicts and quarantined submissions, then accepts or returns
the inspection. Acceptance is recorded in the audit trail and starts the purge
timer on the server side.
Edge flows
App killed or tablet rebooted mid-form: the inspector reopens to exactly the
last saved state.
Grace period expired while offline: the app locks. Unsynced data stays
encrypted on the device, and re-authentication starts the sync.
Storage full: the app blocks further photo capture with an explanation.
Form entry keeps working.
Returned for rework after purge: the inspection downloads again as a new
offline copy.
5. Local data model (SQLite)
SQLite changes from read-through cache to local source of truth for drafts.
Server-derived data and local changes are kept apart, so the device can always
rebuild from the server and replay its own changes on top.
inspection
id TEXT PK -- server id
template_id TEXT
template_version INTEGER -- pinned at download (A9)
server_version INTEGER -- last integer version seen from server
server_snapshot BLOB -- last server copy (JSON, encrypted); base for 3-way compare
local_state TEXT -- downloaded | in_progress | submitted_pending_sync
-- | synced_awaiting_receipt | purged | quarantined
downloaded_at, last_synced_at INTEGER (device ms) + server timestamps where known
purge_after INTEGER -- backstop deadline (A1/D3)
field_value
inspection_id, field_path TEXT -- stable path, e.g. "roof.condition" or "defects[<row_uuid>].severity"
value BLOB -- encrypted JSON
base_value BLOB -- server value this edit was based on (for merge)
dirty INTEGER -- has un-acknowledged local edit
updated_at_device INTEGER
PK (inspection_id, field_path)
repeat_row
inspection_id, section, row_id TEXT (client UUIDv4/v7), deleted INTEGER (tombstone), sort_key
outbox_op -- append-only, ordered
op_id TEXT PK -- client UUID; idempotency key
inspection_id TEXT
seq INTEGER -- per-inspection monotonic
kind TEXT -- set_field | add_row | delete_row | attach | detach | sign | submit
payload BLOB -- encrypted; includes base_value
created_at_device INTEGER
state TEXT -- pending | in_flight | acked | rejected
attempts, last_error
attachment
id TEXT PK -- client UUID
inspection_id, field_path, caption
file_path TEXT -- app-private, encrypted file
sha256, byte_size, mime, captured_at_device
upload_state TEXT -- pending | uploading | uploaded | verified
upload_session_id TEXT, bytes_confirmed INTEGER
signature
inspection_id PK, image_attachment_id, strokes BLOB, content_hash, signed_at_device, signer_name
sync_meta -- per device
device_id, last_pull_cursor, last_successful_sync, auth_grace_expires_at
local_audit -- device-side events not implied by ops (opened, viewed, purge executed)
id, inspection_id, event, at_device, uploaded INTEGER
Rules:
Each UI edit makes one transaction that writes field_value and appends an
outbox_op. If either write fails, the edit is not acknowledged.
Repeated edits to the same field may be coalesced in the outbox while still
pending, which keeps backlogs small. The audit then records the final value
plus device timestamps. Whether each keystroke-level intermediate value must be
audited is a compliance question. The recommendation is no: audit per committed
field value (on blur, or after a debounce).
WAL mode, synchronous=FULL for outbox writes, and foreign keys on.
Customer data columns and files are encrypted (D5). sync_meta, and stubs with
no customer data, do not need to be.
6. Sync protocol
6.1 Pull (server to device)
GET /me/assignments?since=<cursor> returns assigned inspection IDs with their
version, state and template_version. The cursor makes the call
incremental.
For each inspection that is new or changed, GET /inspections/{id} (existing
endpoint) returns the full record, and it is stored as server_snapshot.
Local fields that are not dirty are refreshed from the snapshot. Dirty
fields keep the local value, and the new snapshot becomes the base for the
compare on the server.
Templates and reference data are fetched by version and are immutable, so they
can be cached indefinitely.
6.2 Push (device to server), recommended option D1-A
POST /inspections/{id}/changes
Idempotency: each op has op_id; the batch has batch_id
Body: {
device_id, base_version, // integer version the device last saw
ops: [ { op_id, seq, kind, field_path, base_value, value, created_at_device }, ... ]
}
Response 200: {
new_version,
results: [ { op_id, status: applied | applied_with_conflict | superseded | quarantined | duplicate,
server_value, conflict_id? } ],
receipt_hash // hash of applied content for this inspection
}
Ordering: each inspection's ops are sent in seq order, and batches are
capped by size. Different inspections sync independently, so one stuck
inspection does not block the others.
Exactly-once effect: the server stores op_id in a dedup table in the same
transaction as the change. A replayed op returns duplicate with the original
result. This covers the "committed, but the response was lost" case, which is
common with intermittent signal.
Merge on the server: for each op the server compares base_value with the
current server value.
Equal: nobody else changed the field. Apply the op. version increases once
per batch.
Different: a concurrent edit. Apply the D2 category rule, write a conflict
record, and return applied_with_conflict or superseded.
The record-level base_version is used for diagnostics and fast paths. The
merge decision is made per field. No backfill of per-field versions is
needed.
Rows:add_row with a client UUID is idempotent. delete_row is a
tombstone. An edit to a row another party deleted resurrects it and is
flagged.
Submit: the submit op is only accepted once every attachment named in the
inspection's manifest has status verified. Otherwise the server returns
pending_attachments with a list of what is missing, and the client retries
after uploading.
Retries: exponential backoff with jitter, capped. Sync is triggered by
connectivity change, app foreground, the OS background-task scheduler and a
manual "Sync now". Only 4xx responses that cannot be retried mark an op
rejected, and a rejected op is surfaced in telemetry and on the web. It is
never deleted.
Receipt: after everything is applied and the attachments are verified, the
server returns a receipt_hash. The device compares it with its own
calculation and only then marks the inspection synced, which makes it
eligible for purge.
6.3 Web supervisor writes
The existing web endpoints stay as they are, but must (a) increment version,
(b) write audit entries through the same audit service, and (c) populate the
field categories used by D2. The web editor shows conflict flags and the
"checked out on device" badge.
6.4 Compatibility
The server supports clients with and without offline support at the same time.
Endpoints are additive and nothing existing is removed.
The client sends X-Client-Capabilities: offline-sync/1. The server can refuse
a protocol version it no longer supports with a clear upgrade error, and never
silently drops ops.
7. Conflict policy (summary of D2, D8, D10)
Situation
Outcome
Visibility
Inspector and supervisor edit different fields
Both applied
Audit only
Same field, and the field is inspector-captured
Inspector value applied; supervisor value kept
Conflict flag on web; audit
Same field, and the field is supervisor-owned
Supervisor value kept (device cannot edit these fields; defensive)
Audit
Inspector edits a repeating row the supervisor deleted
Row restored with the inspector's edit
Conflict flag
Inspection reassigned or cancelled while offline
Changes accepted into quarantined
Supervisor queue; device notice
Inspection already accepted when the device syncs
Changes accepted into quarantined, and the accepted content is not modified
Supervisor queue; audit
Template retired while offline
Stored against the pinned template version
Flag if the template can no longer be rendered
Edit after signing
Allowed with a prompt to collect the signature again; signature content hash mismatch recorded
Web shows "signed content differs"
Principle: the system never silently discards a value. Every losing value is
kept in the audit and in the conflict record.
8. Attachments and signature
Capture: photos are taken in the app's camera, compressed (D6), encrypted,
and written to app-private storage. The SHA-256 hash is computed over
the compressed image before encryption, so the server can check the same bytes
it receives. The attachment row and its attach op are written
in one transaction after the file is fsynced.
Limit: 40 photos per inspection, enforced on both client and server.
Storage is checked before the inspection starts (A7).
Upload: photos go through a separate queue from the field ops, so a few
kilobytes of form data are never blocked behind megabytes of photos. Uploads
are resumable and chunked:
POST /attachments/sessions creates a session and is idempotent by attachment
ID. PUT chunks with offsets. POST .../complete makes the server verify the
hash, and the attachment becomes verified.
Order: form ops are uploaded first, then photos in capture order. Upload
runs in the OS background-transfer service where the platform allows it.
Integrity: the server rejects a hash mismatch, and the client uploads the
file again from the start.
Signature: stored as vector strokes and a rendered PNG attachment, plus the
content_hash (D7). The signer's name is a captured field.
Server storage: objects are stored under non-guessable keys. Access goes
through the API's authorisation, never through public URLs.
9. Migration and backfill
Mobile
A forward-only SQLite migration adds the new tables. The old read-through
cache tables are dropped and fetched again (A12). Before dropping them, the
migration scans for any user-authored draft data and turns it into
field_value rows plus outbox ops if it finds any.
The migration runs on first launch after the update, inside a transaction, and
is tested from every app version still installed (MDM inventory).
Downgrade guard: an older app build would not understand the new schema.
The rollback runbook (section 13) requires the outbox to be drained before any
downgrade is pushed through MDM. The new build writes a schema_version
marker, and a downgrade-safe hotfix build (if one is needed) must refuse to
drop a database whose outbox is not empty.
Server
Additive schema: an op_dedup table (op_id, inspection_id, result, and a
TTL well beyond the maximum offline period), a conflict table, an
attachment_upload_session table, the device registry and purge receipts,
field category metadata on templates, and a quarantined record state.
Audit backfill: existing inspections get a single "baseline" audit event
that captures their current state and version, so each audit chain has a
start. Historical changes are not reconstructed unless compliance requires
it. That is an open question for M0 and must not be guessed.
Field categories: templates are annotated as inspector-captured or
supervisor-owned. Where a field is ambiguous, the default is
inspector-captured, and product reviews the list.
Repeating rows: if the server stores repeating sections as positional
arrays (A2), existing records get stable row IDs in a one-off backfill. The
backfill is idempotent and verified by comparing record hashes before and
after, apart from the ID fields.
Existing web clients are not affected. All migrations can be deployed before
any client uses them.
10. Security and data retention
Encryption at rest: database and photos are encrypted with keys in the
hardware-backed Keystore/Keychain (D5), in addition to MDM full-disk
encryption.
Authentication offline: a grace period (D4) and an app lock that requires a
device unlock or app PIN to resume. The refresh credential is bound to the
device (A4).
Transport: TLS only, with certificate pinning if the organisation's policy
allows it and MDM proxies do not break it.
Authorisation: the server re-checks on every sync that the user may still
write to this inspection. Changes from a user who is no longer assigned go to
quarantine (D8) and are not applied.
Purge (D3 and A1):
Primary trigger: server receipt confirmed. The device deletes field values,
snapshots, photos and the signature for that inspection, then destroys the
per-inspection data key (crypto-erase) and keeps only a stub.
Backstop: a purge_after deadline on every downloaded inspection. It is set
from the server's acceptance time when known. Otherwise it is set to a
maximum retention since the last server contact that compliance agrees to.
The backstop never deletes unsynced data on its own. If the deadline and
the unsynced state collide, the device raises an alert and the case is
escalated. Compliance has to decide how that edge case is handled.
The device sends a purge receipt (inspection ID, time purged) on its next
sync. The server reports every accepted inspection without a purge receipt
after 24 hours, and MDM remote wipe is the last resort.
Checks run on app start, on sync, and in a scheduled background task.
Leakage controls: app data excluded from OS and cloud backups; no photos in
the shared gallery; screenshots and the recent-apps preview blocked on
inspection screens; no customer data in logs, crash reports or analytics (IDs
only); clipboard use limited.
Audit: every server change records the actor, device ID, op_id, device
time, server time, old value, new value, the conflict outcome, and the base
version. Device-only events (open, purge) are uploaded from local_audit.
Audit records are append-only.
Threat review: a security review of the offline design is required before
pilot (M5 exit).
11. Observability
Telemetry contains no customer data and is keyed by device ID, inspection ID and
app version.
Client metrics: outbox depth and age of the oldest pending op; attachment
queue depth and bytes; how long sync takes after connectivity returns; sync
attempts and outcomes by error class; conflict counts by field category; local
validation failures; database size and free storage; migration success or
failure; time from submit to server receipt; purge executed or overdue; crashes
tied to the sync state.
Server metrics: changes endpoint latency and error rate; duplicate op rate
(this shows how often retries happen); conflict and quarantine rates; attachment
hash mismatches; inspections accepted without a purge receipt after 24 hours
(compliance alert); age of the oldest unacknowledged conflict.
Integrity canary: each client reports how many ops it has generated and
acknowledged per inspection. The server compares that with the ops it has
applied. Any gap that lasts longer than a threshold triggers an alert, because it
points to lost data.
Dashboards and alerts: a pilot dashboard per cohort. Alerts go to the owning
team for the compliance purge breach, integrity canary gaps, rejected ops, and
spikes in the changes endpoint error rate.
Support tooling: a support view showing a device's sync state (queue depths,
last error) with no customer content. A "Send diagnostics" action in the app
uploads the same data.
12. Staged rollout
Offline mode is gated by a server-side feature flag per user and device. There
are two separate switches: offline capture (downloading and editing offline)
and offline sync (draining the outbox). The sync switch is never turned off
for devices that hold unsynced data.
Stage
Cohort
Entry criteria
Exit criteria (to advance)
S0 Internal
Engineering/QA devices, test data only
M1 to M6 done
All test plan suites green; offline chaos run with zero lost ops
S1 Dogfood
Internal staff doing mock inspections in real basements, on the pilot device model
S0 exit; security review passed
Zero integrity canary gaps; purge receipts for 100% of accepted inspections within 24h
S2 Pilot
One field team (D11)
S1 exit; compliance sign-off on A1/D3; support runbook ready
Across a pilot period agreed with product and compliance: no confirmed data loss; zero purge breaches; conflict rate and quarantine rate within the thresholds agreed in M0; inspector satisfaction ≥ baseline; no severity-1 incidents open
S3 Expansion
Further teams in increments, lowest-connectivity teams first
S2 exit; any pilot fixes shipped
Same metrics hold at each increment; server capacity headroom confirmed
S4 GA
All inspectors
S3 exit
Old online-only code path scheduled for removal after a stable period
The thresholds are placeholders to be set in M0 using the baseline measurements
(for example today's rate of lost notes from support tickets). This plan does not
invent them.
13. Rollback
Kill switch: turning off offline capture stops new offline downloads
and edits. Devices go back to online-only editing, but keep draining their
existing outbox and photo queue. Nothing on the device is deleted by a flag
change.
Server: the new endpoints are additive. Rolling back server code means
disabling the changes endpoint only after telemetry shows every outbox is
empty. Records written through the endpoint are ordinary records with
ordinary versions, so the web and old clients read them normally.
Client downgrade: only through MDM, and only after the outbox is drained
(see section 9). If a client defect blocks draining, ship a forward fix. Do
not downgrade. Before pilot, prepare an emergency "export unsynced data"
support path: an encrypted bundle uploaded through a separate endpoint.
Data repair: because the audit is append-only and records old and new
values, a faulty merge rule can be corrected by replaying the audit. A replay
tool is built in M3 and tested before pilot.
Decision rights: the runbook names who can pull the kill switch, which
metric levels trigger it (integrity canary gap, purge breach, crash spike), and
how pilot users are told.
14. Testing strategy
Layer
What
Notes
Unit
Merge function (every row of section 7), outbox coalescing, state machine transitions, purge rules, content hashing
Property-based tests: arbitrary interleavings of device and web edits never lose a value (it is either applied or recorded as superseded)
Server integration
Idempotency (replay of every op and batch), concurrent web and device writes, quarantine paths, submit gated on attachments, audit completeness
Includes the "commit succeeded, response dropped" case
Client integration
Real SQLite + encrypted storage with a fake server
Kill the app at every step of the sync loop and assert on the recovered state
40 photos on the lowest-spec tablet; backlog of several inspections; battery use during background sync; database size
Encryption overhead measured (D5)
Clock and time
Device clock skewed forwards and backwards; timezone change; long offline period beyond the grace period
Ordering must never depend on device time
Migration
Upgrade from every installed version (MDM inventory), with and without cached data; server backfill idempotency
Test downgrade guard behaviour
Security
Forensic inspection of the device after purge (no customer data in the DB, files, WAL or logs); backup exclusion; screenshot blocking; token expiry offline
External or internal security review
Compliance
End-to-end acceptance followed by a purge receipt in under 24h, including a device that goes offline right after submitting; audit trail completeness report
Compliance signs off on test evidence
Field
Dogfood and pilot in real basements; structured feedback from inspectors
Compare lost-work reports with the baseline
15. Milestones
Milestones are ordered by dependency, not by date. Each milestone ends with a
demo and a check against its acceptance criteria.
Validate A1–A12; decide D1–D11; baseline metrics (lost-work tickets, concurrent-edit frequency); compliance interpretation in writing; team sizing
–
Every decision recorded with an owner; A1, A2 and A3 confirmed or the plan revised; baseline numbers captured; S2 thresholds agreed
M1 Durable local drafts
New SQLite schema, encryption (D5), migration from the cache, transactional autosave, outbox recording (no push yet; while online the existing save path is still used)
M0
Killing the app or rebooting at any point loses zero committed edits (automated test); migration passes from all installed versions; encryption overhead within the agreed budget on the lowest-spec device. This already fixes the "lost notes" symptom when online and can ship behind a flag.
M2 Server changes API and audit
/changes endpoint, op dedup, field-level merge with D2 categories, conflict and quarantine records, audit service, baseline audit backfill, row ID backfill if needed, web conflict flags
M0
Idempotency suite passes (every op replayed N times gives the same result); audit entries for 100% of writes from web and device; existing web and client regression suite green; backfill verified by hashes
M3 Offline sync engine
Pull with cursor, push outbox, retries and backoff, auto-download window (D9), sync status UI, "checked out" badge (D10), audit replay tool
M1, M2
Chaos suite: zero lost or duplicated ops over a long randomised run; conflicts resolved according to section 7 in end-to-end tests; offline open, edit and submit works in airplane mode
M4 Attachments and signature
Camera capture, compression (D6), encrypted storage, resumable uploads, submit gated on attachments, signature with content hash (D7)
M3
40 photos captured and uploaded over a flaky link with no hash mismatches left unresolved; upload resumes after app kill; submit is blocked until all attachments are verified; signature hash mismatch shown after an edit
M5 Security and purge
Offline auth grace period (D4), app lock, purge on receipt plus backstop, crypto-erase, purge receipts, compliance report, leakage controls
M1, M2 (and M4 for photo purge)
Forensic test finds no customer data after purge; purge receipts arrive for a device that went offline right after submitting; security review passed; compliance signs off on the purge design
Every metric in section 11 visible on a dashboard for test devices; alerts fire in a staged drill; the rollback drill completes without data loss
M7 Internal and dogfood (S0–S1)
Run S0 and S1
M6
S0 and S1 exit criteria met
M8 Pilot (S2)
One field team, with weekly reviews of the metrics and feedback
M7, D11
S2 exit criteria met; go/no-go recorded
M9 Expansion and GA (S3–S4)
Incremental rollout, then plan removal of the old path
M8
S3 and S4 exit criteria met
M1 and M2 can proceed in parallel after M0. M5 can start once M1 and M2 exist,
running in parallel with M3 and M4.
16. Open questions to take into M0
What exactly is "local customer data" under the regulation, and does it cover
device logs, OS backups or stubs holding only an ID? (A1)
Must the audit history include intermediate values before commit, and must it
reconstruct historical changes from before launch? (sections 5 and 9)
What is the maximum time a device may keep unsynced customer data, and what
happens when that conflicts with the purge rule? (section 10)
Which template fields are supervisor-owned? (D2)
What photo resolution is acceptable as evidence? (D6)
How long should the offline auth grace period be? (D4)
opus-5.5/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.75 / 10graded blind as submission F
Same family as D (purge every registered copy before acceptance), but better organized. It has an explicit replica registry, epochs, fencing and a broad verification matrix. The risk reasoning on retention, clocks and copies is excellent. Product judgment suffers: acceptance can be delayed indefinitely for a missing tablet, aggregate-version conflicts force review even for disjoint edits, and any content change requires the signer again.
Strengths
Section 8 gives a stepwise acceptance and deletion protocol with fail-closed invariants
M1 and M2 run in parallel after M0; each milestone has concrete exit criteria
Separate flags for prepare, capture, ingestion and acceptance; rollback preserves work
Migration addresses legacy plaintext copies and unknown cached copies
Weaknesses
Over-engineered for a pilot (replica epochs, quiesce and fence, revocable media URLs)
A lost tablet blocks acceptance indefinitely
Conservative aggregate conflicts and the re-signing burden
Evidence the grader checked
§8 'A lost tablet blocks acceptance until a verified wipe is possible'
§7 'Recommend conservative aggregate-version conflicts for the pilot, even for disjoint edits'
§13 milestone table
Objective checks
Files
gpt-6-astra/06-product-planning/RESPONSE.md
Three riskiest assumptions
PLAN.md covers every requested planning area and includes decisions with recommendations, milestones with dependencies and acceptance criteria, rollout/rollback, and a verification matrix. Its technical choices are proposals based only on BRIEF.md; implementation and regulatory approval have not occurred.
Acceptance can wait for every local copy to be removed. The 24-hour deadline starts when the server accepts an inspection. An unreachable or powered-off tablet cannot reliably learn that event or execute a cleanup timer. The recommended pilot therefore freezes a complete submitted revision, reconciles all pending work, purges every registered tablet copy, and only then accepts it. This assumes product accepts potentially indefinite delays for a missing tablet, every copy can be tracked, and compliance accepts the verified erasure method. A lost acknowledgement must delay acceptance rather than bypass cleanup. Validate the workflow with product/field operations and demonstrate key destruction, file/cache cleanup, and backup exclusions on real managed tablets with security/compliance. If immediate acceptance or literal physical erasure is required, this recommendation is insufficient: redesign and prove the guarantee before shipping. See D1, D2, D9 and section 8.
All server writers can adopt one transactional version, retry, and audit contract. Existing REST APIs with integer record versions do not establish aggregate concurrency checks, idempotency, or atomic audit history. If the web app, a legacy endpoint, or an administrative job bypasses the new contract, it can silently overwrite offline work, duplicate an action after a timeout, invalidate signed content unnoticed, or bypass the acceptance fence. Inventory every writer and prove that the version check, mutation, idempotency receipt, audit event, and change notification commit together. Test concurrent web/tablet writes and a lost success response before integrating field sync. Require compatible clients for pilot inspections; if legacy paths cannot be fenced, restrict the pilot or delay it. See D3 and milestones M2–M3.
A prepared inspection is a complete, usable offline work package. The brief does not establish how long tablets stay disconnected, whether every form's lookups/validation are local, how large 40 usable photos are, or whether a signer can return after a supervisor edit. The plan assumes work can be downloaded before entering a basement, fits reserved tablet storage, stays within an approved offline authorization window, and can be re-signed if reviewed content changes. Failure means inspectors can start work they cannot complete or submit, even if synchronization is technically correct. Walk through a real field/review cycle using supported forms and the lowest-capacity managed tablet; measure media sizes, storage/battery use, offline duration, conflict frequency, and re-signing feasibility. Set explicit package limits and pilot eligibility from that evidence. Exclude unsupported forms and reconsider the signature/review workflow if those assumptions fail. See D4–D8 and milestones M0–M1/M6.
gpt-6-astra/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
gpt-6-astra/06-product-planning/PLAN.md
Offline inspection plan
This is a proposed implementation plan, not an approved specification. BRIEF.md is the only product input. Requirements below come from that brief; policy choices, budgets, roles, and technical contracts are recommendations that need validation. No existing code, infrastructure, or regulatory interpretation has been verified.
1. Outcome and constraints
Inspectors on company-managed tablets must open assigned inspections, fill forms, take up to 40 photos, and collect a signature without a connection.
Notes must survive a lost connection, including loss immediately after opening an inspection. SQLite currently acts as a read-through cache and must become durable working storage.
Intermittent connectivity must make progress without duplicating or losing work.
Supervisors can edit the same inspection through the web; concurrent changes must be detected and resolved explicitly.
The system must maintain an audit history and remove local customer data within 24 hours after acceptance.
The existing integration is REST with integer record versions. Delivery starts with a pilot and excludes a full collaborative editor.
Success means that every acknowledged local save survives restart, a complete offline inspection reaches the server exactly once in effect despite retries, competing edits remain recoverable, and acceptance cannot create an unfulfilled local deletion obligation. Pilot thresholds and evidence appear below.
2. Assumptions and decisions required before implementation
The decision owners are proposed responsibilities, not staffing commitments. All decisions are open. Record the owner, chosen policy, and evidence before dependent implementation; a recommendation here does not count as approval.
ID
Assumption or decision
Recommendation and rationale
Required validation / owner
D1
What does “accepted” mean, and can acceptance wait for offline tablets?
Treat acceptance as a terminal server transition performed by a supervisor. Gate it on confirmed purge from every registered local copy, as described in section 8. Show “awaiting tablet cleanup” before acceptance.
Product and compliance must approve delayed acceptance, including a missing tablet. If immediate acceptance is mandatory, redesign the retention mechanism before implementation.
D2
What constitutes removal of local customer data?
Use per-inspection encryption keys, destroy them, and remove all associated database rows, files, thumbnails, temporary data, and UI caches. Exclude customer data and keys from backups.
Security/compliance must approve the erasure definition, whether its scope is inspection-specific or customer-wide across other active inspections, and managed tablet/OS capabilities. Do not claim that deleting a SQLite row proves physical erasure on flash. The proposed key model assumes inspection-specific scope.
D3
Can every writer participate in one concurrency and audit contract?
Require conditional writes using one aggregate inspection version, atomic idempotency receipts and audit events, and the same acceptance fence for mobile, web, legacy endpoints, and administrative jobs.
Backend/web owners inventory all write paths and prove the contract. Child-record versions alone may be insufficient.
D4
Can inspectors prepare work while connected? How long must offline access last?
Require explicit preparation of assigned inspections before entering a basement. Initially allow one tablet assignment per inspector, but track all copies and replacements. Select a finite offline authorization duration from field needs and security policy.
Product/field operations validate a real shift, reassignment, device replacement, and maximum expected offline interval. Uncached work cannot first download without a connection.
D5
Are existing forms suitable for offline use?
Pin the complete form/schema and validation rules when preparing an inspection. Finish against that revision; do not silently replace it during work.
Mobile/backend owners inventory remote lookups, conditional fields, and mandatory schema changes. If a form cannot run offline, exclude it explicitly from the pilot.
D6
What does a signature attest to, and who resolves conflicts?
Bind a signature to the form content and ordered photo manifest. Any change to that content requires a new signature. Prefer resolution by the inspector where authorized; route restricted fields to a supervisor.
Product/compliance approve attestation, signer presence, resolver permissions, and how re-signing occurs after a supervisor edit.
D7
What are attachment and storage limits?
Interpret the limit as 40 photos per inspection plus a separate signature. Choose maximum encoded photo bytes, dimensions, signature bytes, and prepared inspection count using real tablet/storage measurements. Reserve space for temporary copies and SQLite journaling.
Product/mobile owners approve image legibility and capacity. Specify whether existing photos count toward 40; recommended: all active photos count.
D8
How do assignment revocation, logout, and offline authorization expiry work?
Stop new offline work when authorization expires, retain protected pending work for authenticated recovery, and recheck permission on every sync. Block account switching while unsynced work needs recovery. A remote revocation cannot be learned while disconnected.
Security/product approve the offline exposure window, recovery role, draft retention policy, and an explicit discard path if needed.
D9
What is already cached on deployed tablets?
Register and migrate every known copy before acceptance enforcement. If legacy copies cannot be enumerated or cleared, use only newly created inspections on a verified, cleaned pilot fleet.
Mobile/operations inventory app versions, OS backups, old photo paths, shared caches, and offline tablets. Unknown legacy copies are a rollout blocker.
D10
Who operates and supports this feature?
Assign responsibility for server sync/audit, mobile storage/purge, web conflict/acceptance UX, and pilot incident response. Set draft, audit, conflict-candidate, and orphan-upload retention policies separately.
Engineering/operations/compliance approve ownership and pilot service thresholds. Server audit retention is not specified in the brief.
Additional engineering assumptions: no OS background execution or device clock is trusted for the deletion guarantee; network requests may duplicate, reorder, or finish without a response reaching the app; a process may die between any two persistence steps. Pilot device count, concurrent active inspections, and traffic are unknown. Size the pilot from measured capacity, not invented scale or headcount.
3. Approach and scope
Approach
Tradeoff
Decision log
Durable SQLite workspaces, an outbox, and REST conditional writes
Fits the current stack; requires explicit replication, attachment, and conflict logic.
Recommended. Makes local persistence and server authority clear.
Replicate a whole database through a new sync platform
Could provide replication primitives, but introduces a platform migration and still needs signature, audit, and acceptance policies.
Defer; no evidence justifies the migration.
Collaborative operation log / CRDT editor
Can merge richer concurrent editing, with substantially more semantics and UX to maintain.
Exclude; the brief rejects a full collaborative editor.
The MVP includes preparing assigned inspections; durable form edits, photos, and signatures; restart recovery; foreground and opportunistic background sync; visible conflicts and retry states; authorized conflict resolution; audit history; and enforced acceptance cleanup. Losing connectivity after opening a prepared inspection must be harmless to saved work.
Exclude offline creation of unassigned inspections, offline acceptance, live cursors, collaborative text merging, videos, and changes to accepted inspections. Offline reassignment and cross-tablet handoff require a later explicit design. Multiple copies still must be detected and handled safely. Broad historical downloads are outside the pilot.
4. Inspector and supervisor flows
Prepare while connected. Authenticate, fetch assigned work, and choose “Make available offline.” The server registers the replica before serving customer data. Download the inspection, pinned form and lookup data, and necessary existing media. Check permissions, storage reservation, schema support, and integrity. Show “Ready offline” only after an atomic installation of the complete package. A partially downloaded package stays visibly unavailable; a newly assigned but uncached inspection shows “Connect to download.”
Open and work. Read and write the local workspace even when connected. Persist each edit and its outbox entry together before showing “Saved on tablet.” A visible “Saving…” state covers uncommitted input; flush on navigation and backgrounding. A storage error must block a success indicator and explain how to recover. Never switch to a network-only write path on reconnect.
Capture evidence. Show photo count and available capacity. A photo is attached only after its durable encrypted file and manifest entry exist. Enforce the 40-photo limit locally and on the server. Collect the signature after the content is ready; display which inspection content it covers. Editing signed content invalidates the current signature and requests re-signing.
Submit offline. Validate the pinned form, photo manifest, and signature locally, then queue submission. Use “Submission pending sync,” distinct from server submission or acceptance. Lock the submitted local revision; an explicit return to editing cancels an unsent submission or creates a new audited revision after reconciliation if it was already sent. Resolve any uncertain in-flight submission first.
Reconnect. Refresh authentication, process control messages and server changes, sync durable attachment bytes, then sync dependent edits and submission. Show counts for pending changes/photos, the last successful server sync, and actionable errors. Connectivity alone is not success. Network flapping should resume work already done.
Resolve conflicts. Preserve local work and show “Needs review,” including base, tablet, and current server values. An authorized user chooses a resolution, validates it, and saves against the current version. Restricted conflicts are escalated to a supervisor through a protected server candidate. If content changed, collect a new signature before submission. Other inspections can continue syncing.
Review and accept on web. Supervisors see submission completeness, conflicts, signature validity, and tablets awaiting cleanup. Their edits use the same conditional write contract and invalidate signatures when relevant. Acceptance begins the coordinated cleanup in section 8. It is not marked accepted until that protocol completes.
App states distinguish downloading, ready offline, saved locally/pending sync, syncing, authentication needed, conflict, submission pending, submitted, awaiting cleanup, and accepted/removed. Avoid a single ambiguous “Saved” label. A removed inspection leaves only a non-customer cleanup receipt; customer details and thumbnails must disappear from lists and previews.
5. Local data and durability
SQLite becomes the authoritative store for tablet work until acknowledged by the server. The server remains authoritative for shared versions, permissions, submission, and acceptance. Never evict a prepared inspection or pending work to refresh a cache.
Local entity
Minimum contents and invariants
Workspace / replica
Opaque inspection handle, tenant/account scope, device installation and replica IDs, grant epoch, lifecycle state, key reference, schema revision, server base version, local revision, offline authorization state. Register on the server before any download; a partial download still counts as a copy.
Base snapshot
Last acknowledged canonical form values, attachment manifest, workflow state, and signature metadata. Kept separately from the working copy for three-way conflict comparison.
Working copy
Current values, pinned template and lookups, local revision, validation state. All customer-bearing fields encrypted with the inspection key; indexing must not expose customer values.
Outbox
Unique operation ID, replica/epoch, local sequence, base version, immutable request payload/hash once dispatched, dependencies, attempt state, and error class. Newer edits cannot mutate an in-flight request.
Attachment
Client UUID, inspection reference, type, encrypted path, plaintext content hash inside the protected manifest, byte length, ordering, upload session/parts, and verified server object reference. Keep local files until cleanup, including after upload.
Signature
Encrypted signature data, declared signer details if required, content hash, capture context, and validity state. Device capture time is evidence with an untrusted-clock label, not server ordering authority.
Conflict / recovery candidate
Base snapshot, local proposal, server snapshot/version, origin operation, and resolution state. Never overwrite this with an incoming server refresh.
Sync checkpoint
Account-scoped change cursor, per-inspection acknowledgement, retry state, and supported protocol version. Persist a cursor only in the transaction that applies its page.
Cleanup receipt
Non-customer opaque replica/epoch, purge-intent ID, completion proof and retry state. Contains no form values, image hashes, names, addresses, or customer-facing inspection identifiers. Remains deliverable after the inspection key is destroyed.
Use transactional local writes and a durability configuration tested against abrupt process/device failure on the supported tablets. A save commits the working change, local revision, and outbox together. Receiving an acknowledgement updates the base snapshot and removes only acknowledged work in one transaction; edits made during the request remain overlaid and queued.
For capture, write an encrypted temporary file, flush it, atomically rename it, and then commit the manifest/outbox transaction. Recovery deletes unreferenced temporary files or offers recovery where a capture was not acknowledged; it flags missing/corrupt referenced files and prevents submission. Filesystem and SQLite changes are not assumed to share a transaction. Maintain a quota/reservation covering 40 photos, signature, existing media, upload buffers, journal growth, and cleanup operations. Never delete pending photos automatically to regain space.
6. REST sync protocol
These are proposed contracts, not claims about existing endpoints. Retain integer record versions, but make the version cover the entire inspection aggregate: form, active attachment references, signature, and business workflow state. Replica/control metadata can have a separate epoch so transport bookkeeping does not change signed content.
Snapshot and pull. A prepare response includes an internally consistent snapshot, aggregate version, template revision, media manifest, replica grant, and a change-feed checkpoint. Registration and acceptance fencing must serialize before bytes are served. An account/device change endpoint returns ordered changes and tombstones with opaque cursors, including assignment revocation and purge intents. Apply each page transactionally. On an expired cursor, fetch a consistent fresh snapshot and reconcile it against the base; never replace the working copy or outbox. The server retains tombstones long enough for active replicas, or requires this full reconciliation.
Push. A conceptual request is PATCH /inspections/{id} with an idempotency key, replica ID/epoch, integer baseVersion, client local sequence, changed fields, finalized attachment references, and signature metadata where applicable. The server authenticates and authorizes every request, validates the grant/fence, checks aggregate version, validates the form and manifest, and atomically commits the new version, audit event, change-feed record, and idempotency receipt. Draft saves allow incomplete forms; enforce types, structure, permissions, and limits on each mutation, reserving required-field completeness for submission. Return the new version and canonical snapshot. There must be no window where the business write commits but its receipt/audit does not.
Use at-least-once delivery with exactly-once business effect per operation ID. An identical retry returns the original outcome; reuse of an ID with a different payload is rejected. Check the stored receipt before returning a version conflict for a retry whose write already succeeded. Scope receipts to tenant, actor/replica, and operation. Retain receipts while their replica epoch can retry; retire them only when stale requests are guaranteed to be rejected. A short fixed deduplication TTL is unsafe for an unbounded outage.
Queue sequencing. Serialize mutations per inspection. Freeze a contiguous prefix of local edits into an immutable batch against its known base version; coalescing unsent scalar edits is allowed only if required audit detail remains intact. Keep subsequent edits in a later local generation. After acknowledgement, advance the baseline and replay only those later edits. Do not just increment an old version for a second offline batch; use the actual acknowledged version. If canonical values differ from the acknowledged proposal, reconcile later edits explicitly. Other inspections may sync independently with bounded concurrency.
Reconciliation and failures. Pull current state first where possible, but still enforce the version on push because another writer can race. Return a version conflict with the current snapshot/version on a stale mutation. Network timeout, connection loss, transient server errors, and rate limits retry with bounded exponential backoff and jitter, honoring retry-after instructions. A timeout is an unknown outcome: query the receipt or resend the identical operation. Authentication expiry pauses for refresh/sign-in; permission loss quarantines protected pending work for authorized recovery. Validation, unsupported-schema, quota, and permanent attachment errors require user action and do not retry forever. A conflict blocks only dependent operations for that inspection.
Submission is a separate idempotent, conditional server transition after all prerequisite mutations and verified attachments are committed. The server checks required fields, 40-photo limit, hashes, absence of unresolved conflicts, and the signature's match to current content. A queued or interrupted upload cannot satisfy submission. An accepted inspection rejects all later mutation requests, including delayed requests from obsolete epochs.
7. Conflict and attachment policies
Conflicts without a collaborative editor
Recommend conservative aggregate-version conflicts for the pilot, even for disjoint edits. This favors understandable review over hidden merging; measure conflict frequency before adding automatic disjoint-field merging. Never use last-writer-wins, client clock ordering, or silent server overwrite.
For example, both clients start at version 12. The supervisor changes a finding and commits version 13. The tablet's notes based on 12 remain local when its request is rejected; they are not lost. Compare base 12, tablet proposal, and server 13. An authorized resolver can retain the new finding and tablet notes in a new request based on 13, with a new operation ID referencing the rejected candidate for audit. Another intervening edit produces a fresh conflict, not forced success. Resolve repeated-row forms by stable row UUIDs and structured fields, not list positions or text diff merging.
Photo additions retain distinct UUIDs; do not automatically union two manifests past 40. A remove-versus-edit, remove-versus-signature, reassignment, form revision change, or workflow transition requires explicit review. Local removal is an audited manifest change and a tombstone until acknowledged; server audit retains the prior reference according to policy. A server-deleted or revoked inspection must not be resurrected by sync. Protect the local candidate until an authorized recovery/discard decision; purge at acceptance remains governed by the copy registry and fence.
Bind the signature to a canonical hash of inspection identity, pinned schema, form values, and ordered active photo UUIDs/content hashes. Exclude transport counters and acceptance bookkeeping. Every customer-content edit, including a supervisor edit or attachment replacement, invalidates the signature. Retain the previous signed revision in the server audit, but do not present it as approval of the new content. Re-signing needs the signer again; this is a workflow assumption, not a cryptographic shortcut.
Attachment transfer and recovery
Use authenticated upload-session creation, resumable chunks/parts, and a finalize endpoint that verifies byte length and whole-file hash. Record the session and confirmed parts locally. On reconnect, ask the server which parts exist; on expired credentials/session, refresh or create a new session for the same attachment UUID and content. Do not recompress a file while retrying it. Compression/rotation happens before the durable manifest/hash is created.
An uploaded blob is initially staging data. Only a verified, authorized reference committed through the versioned inspection mutation makes it active. A finalization or linking retry cannot create duplicate photos. Uploads must not write the tablet camera roll, shared photo library, unencrypted thumbnail cache, or general-purpose temporary directory. Capture app-private encrypted files, including signature images; strip unnecessary location metadata subject to the evidence policy selected in D7.
Garbage-collect unreferenced server staging uploads after a policy-approved grace period that accommodates active resumable sessions. If a session expires, retain local bytes and re-upload. Local orphan cleanup must distinguish files never committed from pending/acknowledged inspection attachments. The pilot records upload bytes and time for a full 40-photo package; quotas and concurrent upload limits are based on those measurements.
8. Acceptance and the 24-hour deletion requirement
The requirement starts at actual server acceptance, not when a tablet next reconnects. “Delete within 24 hours of the next sync” does not satisfy it. A disconnected or powered-off tablet cannot receive acceptance, and ordinary mobile background jobs cannot guarantee a wall-clock deletion deadline. No such guarantee is assumed.
Recommended pilot policy: purge all local copies before acceptance. This changes availability of acceptance, so D1 and D2 are release blockers. The protocol is:
Maintain a server registry of every tablet replica and download grant, including partial downloads, replacement devices, shared-list/cache copies, and any customer-bearing export allowed by the app. Registration precedes serving bytes. A grant is never considered gone merely because it expired or a tablet has not checked in.
A supervisor requests acceptance of a complete submitted version. In a serialized transaction, create an acceptance intent with that candidate version and freeze web/mobile content writes, new downloads, and new grants for the inspection. Display “Awaiting tablet cleanup,” not “Accepted.” The server snapshot, all evidence bytes, audit entries, and valid signature must already be durable.
Ask every registered replica to quiesce. Each tablet transactionally blocks edits, settles unknown operation outcomes, and verifies an empty pending queue and an exact match to the candidate. If any copy has unsynced work, unresolved conflicts, or a different signed revision, stop preparation before deletion; release the fence, reconcile and review, then start a new intent. Never discard pending work to make acceptance pass. A disconnected copy leaves preparation pending.
Once every copy is ready, durably issue epoch-scoped purge intents and invalidate all data-download capabilities for those copies. The fence covers snapshots, change-feed payloads, assignment lists, and customer-bearing idempotency response replays as well as media downloads; receipt-status queries can return an opaque outcome. Wait for active transfers to finish and be accounted for, or revoke them. Long-lived unrevocable media URLs cannot be used here: previously issued URLs could recreate a copy after deletion. Only after this barrier can tablets purge.
Each tablet durably closes the replica epoch, stops/drains its workers, destroys the inspection key, deletes rows/files/temporary media and lookup copies, clears previews and in-memory customer data, and verifies the cleanup inventory. Every asynchronous callback checks the epoch before persisting or displaying a response; a delayed response cannot recreate a workspace or generate a replacement key. The tablet durably records a non-customer receipt and retries delivery if disconnected. On restart, finish any interrupted cleanup before exposing the inspection. A lost acknowledgement delays acceptance; it never causes premature acceptance. After purge starts, keep the candidate frozen and retry forward rather than silently cancelling the intent and reissuing downloads.
After all required purge receipts or verified managed-device wipe evidence are recorded, the server conditionally marks the same candidate accepted and records server accepted_at in the audit. Accepted content cannot be downloaded again to a tablet. Customer-bearing notifications or list rows must not reintroduce it. The acceptance record references the replica set and cleanup evidence used for the decision.
This establishes, under the approved erasure definition and complete-copy registry, that no readable local customer copy exists at acceptance. It is stricter than the 24-hour limit. A cleanup receipt must prove completion, not merely receipt of a delete command. A lost tablet blocks acceptance until a verified wipe is possible; “remove device from registry” is not an override. Operations may escalate the blocked workflow but cannot bypass the invariant.
If product requires acceptance while a copy is unreachable, evaluate enforced per-inspection access leases and hardware-backed cryptographic expiry, with acceptance-aware deadlines and a strategy for pending edits. A periodic timer, device clock, MDM command that has not executed, or an expiring token alone does not remove stored data. Powered-off expiry and legal acceptance of crypto-erasure require proof. Do not ship an immediate-acceptance variant on the assumption that reconnect or background work will happen within 24 hours.
9. Security and audit
Use TLS, managed-device enrollment, device-bound credentials, OS secure key storage, and per-inspection data encryption. The database journal, shared lookup/cache data, thumbnails, crash reports, exports, clipboard behavior, notifications, app-switcher snapshots, and backups are part of the data inventory. A shared database encryption key alone cannot erase one inspection. Avoid plaintext customer data in indexed columns and migrate existing plaintext remnants under D2's approved process. OS-level screenshots or backups must be disabled/controlled by the managed-device policy if they can create untracked copies.
Require online authentication for provisioning and grant issuance. Permit offline unlock only on the enrolled device for the approved authorization window; do not persist passwords. Validate a platform-backed elapsed-time mechanism for that window; an editable wall clock is insufficient. If elapsed time cannot be established after a reboot or clock anomaly, require reauthentication while retaining pending work, and validate this availability tradeoff in D4/D8. After expiry, lock access and preserve encrypted pending work for reauthentication. Reassignment/revocation takes effect on the server immediately and on an offline tablet only when it learns of it or the offline window ends. Document this exposure in D8. Logout and account switching cannot expose another account's cached work; do not silently wipe unacknowledged edits. MDM remote wipe is a recovery control whose execution must be verified, not an offline timing guarantee.
Create an append-only, access-controlled server audit for accepted edits, submission, conflict resolution, signature capture/invalidation, assignment changes, acceptance intents, purge receipts, and acceptance. Include actor/role, device/replica, operation ID, base/result version, changed fields or protected before/after content, server receipt time, and any reported local capture time marked untrusted. Preserve enough evidence to reconstruct accepted content and who changed it. Audit storage is protected and monitored separately from operational logs; administrators cannot silently rewrite history. Define retention and any tamper-evidence requirements with compliance in D10.
The tablet keeps its own encrypted pending-event journal until upload so offline actions are not omitted. Server acknowledgement means the audit and business update committed together. After cleanup, audit history remains on the server; customer-bearing local audit entries must be erased with the rest of the inspection. Telemetry uses opaque correlation IDs and error categories, never form text, signer data, photos, access tokens, or raw request bodies.
10. Migration and backfill
Inventory the actual SQLite schema, cache eviction, current save behavior, form definitions, photo locations, API writers, and device/app versions. Identify whether an existing local change can be distinguished from a server cache value. Do not call ambiguous content disposable or manufacture a historical edit history.
Deploy additive server support first: aggregate versions, transactional receipts/audit, grants/replicas, change cursors/tombstones, resumable uploads, conflict candidates, and acceptance intents. Update every web/legacy write path and require a compatible client for pilot inspections. Fail closed on acceptance if a writer or copy is unaccounted for.
Backfill stable IDs for existing form rows/attachments and an initial aggregate version from a consistent server snapshot. Coordinate backfill with live writes using locking or a change watermark and catch-up pass. Add an audit baseline explicitly labeled “migration snapshot”; it is not evidence of unrecorded historical changes. Validate counts, relationships, and attachment hashes before enabling sync.
Upgrade mobile storage with versioned, restartable migrations. Check space first. Recover identifiable drafts into a protected outbox; quarantine ambiguous cache content for review. Copy and verify into the encrypted format, record a completion marker, then remove the old plaintext store and sidecar files using the approved erasure procedure. Any temporary recovery copy is encrypted and included in cleanup tracking. Disk-full or interruption leaves recoverable state and no false “Ready offline” status.
Register existing pilot copies and refresh consistent snapshots without overwriting recovered edits. Existing accepted customer data gets an immediate cleanup sweep; lack of a known acceptance time does not justify retaining it. Already missed deletion deadlines are surfaced as incidents, not described as fixed retroactively. Old disconnected clients remain a blocker until upgraded/cleaned or verifiably wiped.
If full legacy assurance is impractical, pilot only new inspections on a verified cleaned device cohort and prevent older clients from opening those inspections. Expand only after the legacy fleet migration is evidenced. Verify app downgrades cannot reopen the old cache or bypass acceptance fencing.
11. Observability and operations
Track oldest pending operation and attachment age, queue size, retry/error classes, time from restored connectivity to submission, bytes retransmitted, conflicts by field/workflow, signature invalidations, storage failures, and local save latency. Split time waiting for connectivity or a user from time failing with healthy connectivity. Avoid using client wall clocks alone for deadline evidence.
Track every acceptance intent's candidate and registered replica set, time awaiting quiescence/purge, receipt delivery lag, attempted post-purge downloads, accepted inspections with missing evidence, and local customer-data remnants found by managed-device checks. With the recommended policy, an accepted inspection missing purge evidence is an invariant failure and triggers an immediate stop of new acceptance. Server state cannot independently prove that a tablet erased bytes; test the device-side evidence mechanism and inspect devices during the pilot.
Provide protected support views for operation/replica IDs, current workflow, error codes, and recovery options. Runbooks cover stuck uploads, expired login, conflicts, missing tablets, storage exhaustion, failed migration, cleanup interrupted by a crash, and restoring server service. Support must never prescribe reinstalling or clearing app storage as the first step for pending work. Assign responders and alert thresholds before pilot launch.
12. Staged rollout and rollback
Stage A — capability verification. Complete policy decisions, real-device durability/erasure experiments, and server writer/copy inventory. Acceptance and deletion guarantees are gates, not deferred polish.
Stage B — controlled end-to-end trial. Use managed test tablets and non-production inspections. Exercise an entire airplane-mode inspection with 40 photos and a signature, concurrent supervisor changes, flapping connectivity, purge, and acceptance. Verify customer-data absence in app storage, previews, and supported backup paths.
Stage C — restricted pilot. Enable explicitly selected compatible tablet/inspector and supervisor cohorts, with eligible form types and prepared-inspection limits. Walk field users through preparation, local versus server save state, conflicts, re-signing, and delayed acceptance. Keep direct support available. Select a representative observation window covering at least one full field/review cycle and missed-connectivity/device-recovery scenarios; set its duration before launch rather than inventing calendar dates here.
Stage D — expansion. Expand cohorts only after an agreed sample/window shows no lost acknowledged edits, duplicate business effects, audit gaps, or retention violations; no unresolved severity-one incident; conflict and re-signing workload acceptable to field users; and approved storage, battery, save-latency, and sync-completion targets met. Quantitative performance targets must be recorded from Stage A/B measurements before Stage C. Continue compatibility restrictions until legacy migration is complete.
Use separate server-side flags for preparing new offline work, new capture capability, sync ingestion, and requesting acceptance. Stopping enrollment or new capture must leave existing work readable/editable as policy permits, sync/recovery available, and cleanup active. Cached flags are not a reliable way to stop a disconnected device immediately.
Rollback preserves work. On a severe defect, halt new offline grants and acceptance requests, restrict the cohort, and retain the compatible sync/receipt/upload APIs for already enrolled clients. If writes themselves are unsafe, pause ingestion with a retryable state while keeping outboxes and attachment bytes intact. Resume with the same operation IDs after repair. Do not roll tablets back to the read-through-cache version or drop migrated columns/keys. Keep acceptance fences, purge intents already issued, audit writes, and receipt processing durable across server rollback; finish started purge/acceptance flows safely. Accepted/purged inspections are not resurrected by rollback. Destructive schema contraction and retirement of protocol versions require proof that no active replica or pending recovery depends on them.
13. Milestones, dependencies, and acceptance criteria
Sequence by verified outcomes rather than dates or unsupported staffing estimates.
Milestone
Dependencies
Deliverable and exit criteria
M0: decisions and feasibility
None
D1–D10 have owners and recorded decisions. Demonstrate the approved erasure approach on supported tablets; inventory every writer and copy source. Agree offline window, eligible forms, attachment budgets, performance targets, pilot sample/window, and operational ownership. Reject or redesign any unprovable retention policy.
M1: durable workspace
M0
Restartable SQLite/encryption migration, offline packages, local save/outbox, media capture, storage accounting, and recovery UI. A full offline form with 40 photos and a signature survives restart; no acknowledged save is lost in crash injection; the 41st photo is rejected clearly.
M2: server contracts and web compatibility
M0
Aggregate-version checks, receipts, audit, change feed, replica registry/fences, and upload service. Every writer honors preconditions. Concurrent conditional writes have one winner; dropped responses replay one effect/audit event. Backfill is reconciled under live-write tests. M1 and M2 can proceed independently after M0.
M3: end-to-end sync and review
M1, M2
Reconnect/retry, attachment finalize/linking, submission, conflicts, and re-signing work end to end. A flapping connection converges without duplicate photos; tablet edits made during sync remain queued; unresolved conflicts block submission and preserve both versions.
M4: acceptance and deletion
M3
Complete copy registration, quiesce/fence, purge receipts, cleanup, and terminal acceptance. An offline or missing replica blocks acceptance; reconnect completes it; interrupted cleanup and lost receipt recover. No accepted inspection is downloadable or has a readable local customer copy under the approved model.
M5: migration and rollback rehearsal
M4
Production-like legacy data/device cohorts migrate without discarding recoverable drafts. Unknown copies prevent acceptance. Rehearse flags, service rollback, recovery, and protocol coexistence with pending work and purge intents. Device inspection verifies no legacy plaintext escape path.
M6: pilot
M5
Stage C cohort completes the agreed field/review sample and observation window; dashboards/runbooks function; field users can prepare, recover conflicts, re-sign, and understand cleanup delays. Meet every Stage D gate before expanding.
M7: broader rollout
M6
Incremental cohort expansion, continued invariant monitoring, completed legacy cleanup, and explicit sign-off against the same gates. Retire old protocols only after active replicas and recoveries are drained.
14. Verification matrix
Use deterministic protocol/storage tests, integration tests with both web and tablet writers, fault injection, and real managed-device field trials. Mocks alone cannot prove filesystem durability, background behavior, backup exclusions, or erasure.
Scenario
Required assertion / evidence
Open while online, lose connection immediately; then cold-launch in airplane mode
A prepared inspection and its full pinned form open; saved notes remain editable. Partial preparation and uncached work show a truthful unavailable state.
Crash before/after local save commit and at every media rename/manifest step
Acknowledged edits and captures survive; incomplete captures have a recoverable/explicit state; no referenced attachment silently disappears.
Exhaust storage during typing, photo capture, upload, and migration
No false save acknowledgement, corrupted workspace, or automatic deletion of pending work; recovery remains possible.
Duplicate/reorder requests; drop a successful server response; restart client/server
Stable operation IDs yield one business mutation/audit event; the queue eventually drains; mismatched payload reuse is rejected.
New local edits arrive during an in-flight batch
An acknowledgement removes only the frozen prefix; later edits remain intact and use the actual next base version.
Tablet and web edit same/disjoint fields, repeat rows, photos, and workflow state
Aggregate conflict is detected; both proposals survive; role checks apply; stale conflict resolution is rejected and can retry.
Limits are consistent; verified bytes are required to submit; retries do not duplicate photos; orphan cleanup cannot delete active references.
Signature captured offline, then tablet/supervisor changes a field or photo
Old signature remains audit evidence but fails validation for changed content; submission requires the correct signature.
Permission revoked, credentials expire, account changes, or clock moves backward/forward
Server authorization remains enforced; pending work stays isolated and recoverable; no device clock can authorize acceptance or defeat cleanup.
Schema update, unsupported template, cursor expiry, tombstone, or deleted inspection
Safe reconcile or explicit block; no outbox overwrite, lost signature context, or deleted-record resurrection.
Second copy, partial download, delayed HTTP callback, replayed receipt payload, active download URL, legacy app, or tablet offline during acceptance
All copies are included or acceptance fails closed; a closed epoch rejects late responses and key recreation; purge cannot be followed by a download recreating customer data.
Tablet has hidden pending work when acceptance preparation begins
Quiescence fails before purge; pending work is reconciled and reviewed; acceptance cannot discard it.
Crash during purge, lose purge acknowledgement, power off tablet, or lose device
Key/data cleanup resumes; receipts retry; no acceptance without all evidence; no timeout removes a replica from the requirement.
Inspect accepted data on device and in backups/logs/previews
Customer data, keys, thumbnails, local audit, and old plaintext artifacts are absent under D2's approved erasure definition; only non-customer receipts remain.
Interrupted upgrade/backfill, mixed clients, feature disable, and service rollback
Data remains recoverable, writes remain conditional/audited, cleanup continues, and old clients cannot bypass the acceptance fence.
Representative lowest-capacity tablet and complete field shift
Meet pre-agreed local save latency, full-package storage, battery, reconnect time, and usability targets with 40-photo inspections. Record device and network conditions.
Release evidence includes audit reconstruction of a submitted/accepted inspection, a complete duplicate-request trace, crash-recovery results, a concurrent-edit/re-signing walkthrough, cleanup evidence for every pilot replica, and a completed rollback rehearsal. Any lost acknowledged edit, accepted record without purge evidence, or silent conflict overwrite blocks expansion.
gpt-6-astra/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.75 / 10graded blind as submission J
Well-reasoned and complete. It uses a per-data-class conflict policy (union, three-way merge with conflict records, server-authoritative fields), preserves answers to orphaned fields, adds a shadow rollout stage, and lets the inspector submit while photos are still pending. Purge reasoning is weaker than the leaders: it measures the deadline against server time once acceptance is learned, but the unknown-acceptance case relies on the local clock and vague 'safe defaults'.
Strengths
Clear data-class conflict table; the same-field rule keeps the inspector's value without blocking anyone
M0–M10 table with concrete acceptance criteria; device and server tracks run in parallel after the contract spec
Shadow stage, and parking the outbox on rollback
D5 (submit not blocked by photo upload) shows good field judgment
Weaknesses
The offline purge path depends on 'the local clock passes a known deadline' and leaves the unknown-acceptance case vague
Some gold-plating (mutual TLS device certificates)
Section order is odd: sync comes before the data model
Evidence the grader checked
§8.3 'If the local clock passes a known deadline, the device purges on its own initiative'
§7 class table and §7.1
§12.1 Stage 1 Shadow
§13 D5
Objective checks
Files
mcode-m3.1-flash/06-product-planning/RESPONSE.md
The three riskiest assumptions
Full reasoning lives in PLAN.md (§2 assumptions, §13 decisions). These three are the ones
where being wrong is expensive, expensive to discover late, or both.
1. The server can be extended additively, and its integer record version is actually enforced
The assumption (A2 / A14 / D2). New sync endpoints and a delta feed can be added to the
existing REST API, and the existing integer record version increments on every accepted
mutation and is enforced — a stale write is rejected with the current record returned, not
silently overwritten.
Why it is risky. This one assumption is the foundation under everything else. The integer
version is the merge base: without a trustworthy version, a three-way merge cannot tell
"changed on both sides" from "changed on one side", and every conflict decision in §7 becomes
guesswork. It is also silently load-bearing — a version field that is incremented but not
enforced is a very ordinary thing to find in a codebase, and the brief's wording ("use REST
with integer record versions") does not distinguish the two. If the answer is "the API is
frozen," the transport layer is not a retrofit at all; it is a separate sync service, and
M1–M4 in the milestone sequence change shape.
Blast radius. Total. Every other milestone depends on it, and so do the idempotency,
delta-feed, and rollback-parking designs — rollback is only genuinely reversible if the server
can keep accepting a parked cohort's ops, which is an additive-endpoint requirement.
Cheapest way to find out. One question to the server owner before any design work: send
two writes at the same base version — what does the second one do? A one-afternoon spike,
answering both halves of the question.
If false. Fall back to a versioned, append-only mutation-log surface beside the existing
API, and treat conflict detection as a new server responsibility with its own milestone. Do
not start M2 (the local outbox) until this is settled — the outbox is cheap to build and
expensive to re-shape.
2. The 24-hour removal rule means what the plan assumes
The assumption (A7 / A8 / D3). The clock starts at server-side acceptance, not at
inspector submission; the device is the obligated party; photos and signature images are
customer data subject to the rule; audit evidence is separate and server-side; and
never-submitted drafts are not customer data under the clock.
Why it is risky. It is a legal determination, not an engineering one, and the plan encodes
it in the storage design. Three sub-questions can each be wrong, and the errors are
irreversible in both directions:
If the clock actually starts at submission, purging on the acceptance-based deadline
destroys unsubmitted inspector work.
If photos and signatures turn out to be exempt as audit evidence, the plan over-purges and
destroys evidence it is required to keep.
If never-submitted drafts do count as customer data, the device is holding customer data
outside the 24-hour window for as long as an inspector is offline — a standing violation, not
a one-off.
The device is the worst possible enforcement point: it may be dark for days, it may never come
back, and the compliance evidence it produces is the only proof the rule was met.
Blast radius. Irreversible, and it reaches the storage layout (tiered retention in §8.3),
the purge scheduler, the signature and photo data classes, and the audit design. Retrofitting
this after a violation is not a code change.
Cheapest way to find out. Get compliance to answer three written questions before M0
closes: what starts the clock, what counts as customer data, and what happens to an
unsubmitted draft. The plan's recommendation is on the record in D3, so compliance is
reacting to a specific proposal rather than an open question.
If false. The local storage design changes (the retention tiers may need to be reversed),
and the app may need to refuse offline drafts in some states rather than merely purging them
— which is a product decision, not an implementation detail.
3. Supervisor edits are field-scoped and infrequent enough that field-level merge is the right design
The assumption (A9 / A12). Supervisors editing from the web make relatively rare,
field-scoped changes; two people are not genuinely co-authoring the same text; and the brief's
"no full collaborative editor" constraint reflects real workflow rather than just budget.
Why it is risky. It is the plan's central bet on conflict handling. The entire design
avoids CRDT/OT and character-level merge on the strength of this assumption, and the §7.1 policy
pushes genuine divergences into a human-resolved queue. If supervisors in practice rewrite
whole sections, add and remove form fields while tablets are dark, or genuinely co-author with
inspectors, then: conflict volume grows well beyond what a badge-and-web-resolution flow can
absorb, the orphaned-answers path (§6.3) becomes routine rather than exceptional, and the
product constraint itself has to be reopened. A conflict-rate spike in pilot is the early
warning, but it arrives only after real inspectors have been through real basements.
Blast radius. Large and rework-shaped. The merge engine (§6, §7) and the supervisor
resolution UX are the two most expensive pieces of the plan; a wrong bet here means either
rebuilding the merge layer or shipping a queue nobody can work.
Cheapest way to find out. Before design work, ask two supervisors to describe their
actual editing pattern on a real inspection, and sample recent server audit history for edits
made to the same inspection within a short window by a supervisor and a device. That sample
is a direct measurement of the assumption, and it is data the company already has.
If false. Two viable paths, both cheaper now than later: constrain the supervisor web client
so that structural edits (field add/remove) are blocked while any device has a pending
unacknowledged edit on that inspection, or re-open the product constraint and fund a proper
merge model. The pilot's conflict-rate gate exists so that if this assumption fails, it fails
on a small cohort rather than at full rollout.
Also worth watching, but not top three
A1 — assignments and form definitions are already cached before the inspector loses signal.
Cheap to mitigate (cache the whole assignment index) and cheap to verify, so the downside is
bounded — but if wrong, the feature cannot start any work offline, which defeats the purpose.
A3 — 40 photos is the real throughput driver. Design is already built around it
(resumable chunked upload, priority queues so photos never starve form data, and allowing
submit with photos outstanding so an inspector is never stranded), so this is a capacity
question rather than a design question.
A10 — the existing SQLite cache is genuinely disposable. If any unshipped offline draft
exists, or users depend on cached text, backfill becomes a workstream that the plan does not
budget for.
mcode-m3.1-flash/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
mcode-m3.1-flash/06-product-planning/PLAN.md
Plan — Offline inspection sync (tablet app)
Input: BRIEF.md only. This plan is derived solely from the brief. Where the brief is
silent, the point is recorded as an assumption (§2) or an open decision (§13) rather
than invented as fact.
No dates, durations, or headcount appear in this document. Sequencing is expressed as
dependency order between milestones (§12) plus measurable gates. Effort estimation requires
knowing the team's composition and is deliberately left out rather than guessed at.
1. Summary
Make the tablet app offline-first for the inspector's working copy and keep the server as
the system of record. Local SQLite stops being a read-through cache and becomes a durable
working store with a transactional outbox. Sync is push-then-pull over additive REST
endpoints built on the existing integer record versions. Conflicts are resolved by data
class — union for append-only data, three-way field merge with a recorded conflict (never
silent loss) for scalar fields, server-authoritative for protected fields. Photos upload over a
separate resumable channel that never blocks form data. Rollout is capability-flagged by
cohort, gated on measured health thresholds, with a rehearsed rollback that does not strand
queued work.
The three load-bearing bets, in order: the server can be extended additively and its integer
version is enforced (A2); the 24-hour removal rule is interpreted the way this plan assumes
(A7/A8); and supervisor edits are field-scoped and infrequent enough that field-level merge is
sufficient (A9/A12). All three are summarised in RESPONSE.md.
2. Assumptions
Each assumption is stated so it can be falsified. Confidence is my judgement, not evidence.
A1–A4 and A7/A8 are the ones that would invalidate large parts of the plan if wrong.
#
Assumption
If false
A1
An inspector syncs at least once in coverage before entering a no-signal area, so assignments and the form definition are already cached and the inspection can be opened offline.
The feature cannot start work offline at all — the exact failure the project exists to fix. Needs a separate assignment-distribution design.
A2
The server can be extended additively (delta feed, idempotency key, conditional mutation on the existing integer version) and that version is enforced, i.e. a stale write is rejected rather than silently overwriting.
The whole transport and merge foundation changes. If the version is not enforced server-side, three-way merge has no trustworthy base and a new sync service is needed.
A3
Up to 40 photos per inspection on a tablet that can hold them, and photos — not form data — are the throughput and storage driver on a weak link.
If real inspections are much larger, the storage/purge/time budgets in §6 and §8 need rework.
A4
Tablets are company-managed with a known MDM baseline (encryption, OS floor, remote wipe, app container excluded from OS backup).
The security model in §9 and the purge guarantee in §8 both weaken; unmanaged or rooted devices may need to be blocked from offline mode entirely.
A5
Normally one inspector on one device works an inspection; two tablets on the same inspection is rare.
The merge rules still hold across devices, but conflict frequency and the signature/status rules need extra cases.
A6
Connectivity returns intermittently but within hours, and inspectors re-enter coverage before the end of the shift.
Devices stay dark for days; staleness UX, supervisor visibility of "in progress", and purge timing all degrade.
A7
The 24-hour removal clock starts at server-side acceptance, not at inspector submission, and the device is the obligated party.
Purging too early destroys work; too late is a violation. Direction of the error is decided by legal, not by engineering.
A8
Photos and signature images are customer data subject to that rule (audit evidence is separate, server-side).
Over-purge (we destroy our own evidence) or under-purge (violation).
A9
Supervisor edits from the web are field-scoped and infrequent, not wholesale form restructuring.
Conflict volume and the orphaned-field problem (§6.3) grow; field-level merge may be the wrong design.
A10
The current SQLite store is a disposable read-through cache with no durable user-authored content.
Backfill of in-flight local work becomes a workstream in its own right (§10).
A11
Audit history is append-only and server-authoritative; a device cannot author a trusted audit entry, only a provisional one.
The audit model in §6.5 and the "unconfirmed audit" label need rework.
A12
A full collaborative editor is genuinely out of scope; supervisors and inspectors are not co-authoring the same text in real time.
Field-level merge plus conflict records becomes untenable, and the product constraint itself must be revisited.
A13
The OS suspends the app when backgrounded, so sync cannot be assumed to run in the background.
If the target platform permits reliable background execution, a background sync trigger becomes available and improves first-upload latency.
A14
The integer version increments on every accepted mutation and is visible to clients, not merely a server-internal ETag.
Version-based rebasing and delta feeds cannot be built on it.
A15
A person may use multiple devices, and a device is enrolled and identifiable server-side (user × device).
Sync identity and audit attribution need a different model.
2.1 Platform note
The brief does not name a platform. This plan is deliberately platform-agnostic. The two
platform-specific consequences called out are: app suspension in the background (A13) and
per-user vs per-device encryption keys if tablets are shared between inspectors (D4).
3. User flow
3.1 First run / after a long gap
Inspector opens the app in coverage. The app syncs: assignment index first, then form
definitions and inspection details, lazily per inspection. A progress surface says
"12 inspections ready for offline use", not a spinner.
The app then states what is available offline and when it was last verified. The inspector
does not have to think about connectivity; they only need to know work will not be lost.
3.2 Core offline loop (basement, no signal)
Open — inspector picks an assigned inspection from the cached list. The header shows
Assigned · not yet synced since <time>. Opening does not require a network round trip.
Fill — every keystroke and toggle is applied locally and appended to the outbox in a
single transaction. A quiet indicator reads Saved on this tablet; there is no
"saving…/offline" flicker, because there is nothing to wait for.
Photo — shutter capture writes the file to encrypted app-private storage first, then the
row. The photo is immediately durable, so an app kill, battery pull, or reboot loses
nothing. A per-inspection counter shows 7 / 40.
Sign — capture stroke or finger signature. The rendered image is bound to a hash of the
inspection content at signing time, so later edits visibly invalidate consent rather than
silently reusing an old signature. The inspector sees the signatory name, timestamp, and
consent text before accepting.
Submit for review — this is an offline action and means "submitted, awaiting
review", not "accepted". The app says so explicitly. The device is not blocked by
photos still waiting to upload; it reports exactly how many are outstanding (see D10).
Leave — nothing to flush. The outbox survives process death, reboot, and app update.
3.3 Reconnection
On app foreground, on network regain, on a foreground timer, and on pull-to-refresh, the app
drains its outbox, then pulls what changed, then reconciles. Progress is visible but never
blocking; the inspector can keep working. When everything is drained the indicator goes
All work synced · <time>.
3.4 Supervisor loop (web)
Supervisors see device state without touching the device: last sync time, pending op count,
conflicts awaiting resolution, and any device that has not checked in. They resolve conflicts
by accepting the authoritative value or pulling in the device's preserved variant, and the
resolution is itself audited.
3.5 After acceptance
The device learns the acceptance, computes its purge deadline from the server timestamp,
and purges local customer payloads within the required window, logging evidence of the purge.
The inspector keeps a minimal, non-customer receipt (ids, hashes, timestamps) that supports
the audit history without holding customer content (§8.3).
4. Scope boundaries
In scope
Offline open, edit, and submit of assigned inspections
Offline photo capture (up to 40 per inspection) with resumable upload
Offline signature capture
Offline request to submit for review
Delta pull, three-way merge, conflict detection, conflict records
Supervisor-side conflict resolution and variant retrieval
Compliance purge with deadline enforcement, including the device-offline-past-deadline case
Per-cohort capability flags, kill switch, and support tooling
A migration path for the existing read-through cache
Out of scope (stated so it is not smuggled in)
Real-time collaborative editing — no CRDT/OT, no character-level merge, no live cursors.
Product has said no and the plan respects that (A12).
Offline authoring of unassigned, archived, or historical inspections
Offline supervisor editing — the web client is assumed online
Media beyond photos (video, audio) — not in the brief
Offline PDF/report generation — the server can produce the report after acceptance
Relying on background execution while suspended (A13); sync is foreground-triggered
PKI-backed digital certificates for signatures (D7)
Any rewrite of the existing REST surface beyond additive endpoints (§5.2)
Analytics that carry customer content of any kind (§11.3)
Deferred, with the trigger that would reopen each
Deferred
Trigger to reopen
Video evidence
Regulatory or product requirement, or observed need in pilot field notes
Offline PDF/report
Field complaint after pilot that inspectors need a local artefact before coverage
Real-time co-editing
Measured concurrent supervisor/inspector edit rate in pilot exceeds the level field-level merge handles cleanly
Offline history browsing
Supervisor request for offline access to accepted inspections (this materially expands the local store and the purge surface)
5. Sync protocol
5.1 Client write path (the one rule that matters most)
Every local mutation goes through a single write path that, in one transaction:
applies the change to the local tables, and
appends an outbox row describing the intended server mutation.
No screen, form, or photo component writes directly to SQLite. This is what makes the
invariant in §5.4 true. A hidden local mutation with no outbox entry would be invisible to
sync and would be permanently lost data.
5.2 Transport
Additive REST on the existing API, reusing the existing integer record version (A2, A14).
Two families:
Push — POST /v2/sync/ops with a batch of op envelopes:
{
"op_id": "uuidv7", // idempotency key, generated once at enqueue
"inspection_id": "…",
"entity": "answer|photo_ref|signature|submit_request|delete_request",
"field": "form.q3.notes", // field ops
"value": { },
"base_version": 47, // last server version the client saw for this record
"client_seq": 128, // device-local monotonic counter
"device_id": "…",
"captured_at": "…", // device clock, advisory only
"form_def_version": 3
}
The server answers each op individually with exactly one of:
applied → { new_version }
duplicate → same result as the original application (safe retry, no side effect)
conflict → { current_record, current_version }
rejected → validation failure, with a machine-readable code
op_id is retained server-side for a window at least as long as a device might plausibly
retry (documented explicitly, not left implicit), so a retry after a lost response is safe.
Pull — a change feed, then detail:
GET /v2/sync/feed?since=<token>&scope=assigned → changed inspections with id,
new_version, changed-field mask, and current form_def_version
GET /v2/inspections/{id}?since_version=<v> → delta if the client's base is usable, full
record if not
If the feed token has expired or the client's base is unusable, the client performs a
defined full re-sync of the assigned set. This is a designed, tested path with its own
progress UI — not an error state.
5.3 Sync cycle
Reachability probe — cheap, short timeout, never blocks the UI. Failure is not an
error state; it just defers the cycle.
Drain outbox in client_seq order, batched by count and by bytes. Independent priority
queues so that form data and signature are never starved behind 40 photos (§6.4). Retry
with backoff; validation rejections are quarantined after repeated failure into a visible
"needs attention" item rather than being retried forever.
On conflict — record it, rebase the op onto the returned version, and retry once. A
second conflict parks the op as a conflict record. No unbounded merge loop.
Pull and merge — fetch changed records, run the merge in §6 against the stored base
(§6.2), persist the new base and version.
The displayed state is always a pure function of (server state, local pending ops). If a
state cannot be explained that way, it is a bug. This is asserted by a test, not just a
comment.
5.5 Triggers
Foreground resume, app launch, network-regain callback, after each local write (debounced),
a foreground interval timer while the app is visible, and explicit user pull-to-refresh. No
trigger assumes a background execution window (A13). Triggers are idempotent — several firing
at once collapse into one cycle.
Merge base: last acknowledged server representation
inspection_id, server_version, state_json
audit_local
Provisional device audit entries
id, inspection_id, event, at, confirmed
sync_meta
Device/user identity, tokens' metadata, counters
device_id, last_success_at, feed_token
The whole database is encrypted at rest (§9).
_synced_state is the reason three-way merge works: without the last acknowledged server
value, "changed on both sides" cannot be distinguished from "only changed on one side". It
also makes the merge path uniform — every record merges the same way.
6.2 Files on disk
Photos and rendered signature images live as files in encrypted app-private storage, not
as BLOBs in SQLite, referenced by relative path from photos/signatures. BLOBs in the
database would bloat every delta read, defeat the row-level merge, and make the purge a
single large rewrite instead of targeted unlinking.
Durability order for a capture is fixed: write file → fsync/close → insert row → insert
outbox entry, all in the order that never leaves a row pointing at a missing file.
6.3 Form schema changes while a device is offline
A supervisor can change a form while a tablet is dark. The feed carries form_def_version.
New required field → applied, shown unfilled, inspection marked incomplete rather than
failing to sync.
Removed field with a local answer → the answer is preserved in an orphaned-answers bucket
and surfaced to the supervisor. It is never silently dropped.
Type change → migrate where unambiguous; otherwise orphan rather than coerce destructively.
A filled field an inspector spent time on is customer work. Discarding it is a data-loss
event even if the field no longer exists.
6.4 Priority and ordering
Two independent orderings keep photos from starving the rest:
Data class — form answers, then submit request, then signature, then photos. Small
critical data always drains first.
Within a class — oldest first by capture time.
This is what makes a slow photo upload harmless rather than blocking.
6.5 Audit entries
Device-side audit entries are written as confirmed = 0 and are labelled provisional until
the server accepts them; the server then owns the authoritative entry (A11). The device never
presents a provisional entry as final.
7. Conflict policy
Conflicts are resolved by data class, not by one global rule. The server is always
authoritative; the device is authoritative only for its own pending ops.
Class
Examples
Policy
Append-only / set-like
Photos, audit entries, additional form rows, forward status transitions
Union merge. Both sides' rows are kept. No conflict is possible and none is reported.
Scalar editable fields
Notes, checklist values, free text
Three-way field merge against _synced_state. Unchanged on one side → take the other. Changed on one side only → take that side. Changed on both sides → conflict record (see below).
Server-authoritative. The device may request a transition via a role-checked op; it may never set the field directly.
Deletions
Inspector deletes a photo after upload
Tombstone op, not a server object removal. The deletion is an auditable event; the audit history records it. Deleting after submission is a regulated event, not an erasure (D12).
Delete vs edit
Supervisor deletes an inspection while a device edits it
Server wins. The device keeps its local copy marked pending removal for the retention window, then purges. Deleted is never resurrected by a stale device op.
7.1 Same-field conflict resolution — the recommendation
Both sides changed the same field since the common base:
The server value stays canonical and is what the device displays.
The device's value is preserved verbatim as a conflict record, server-side and visible to
both the inspector and the supervisor.
The inspector sees a non-blocking badge: "Your note differs from the supervisor's update.
Your version is saved. Review in the web app." It never blocks their next inspection.
The supervisor resolves by accepting the authoritative value or pulling in the device variant.
The resolution is itself audited.
Why not last-writer-wins: the device's clock is not trustworthy, and silent LWW in a
basement is indistinguishable from the data loss this project exists to stop. Preserving the
value costs storage and adds a resolution queue, and it is the only option that satisfies both
the audit requirement and the user's actual need.
Why not device-wins: supervisors are accountable for record accuracy, and a field can be
marked complete or compliant by them. Letting an offline device overwrite that is worse than
making a human look at a badge.
The honest cost: this creates a queue somebody has to work. That is a real operational
expense, and the pilot must measure its size (§11).
7.2 Signature and status special cases
Two signatures captured on different devices while offline → keep the earliersigned_at, flag for supervisor resolution, retain the later as a conflict record. The
inspection is not silently re-signed.
A device cannot sign an inspection the server has already accepted; the op is rejected
server-side and the device is told why.
Status transitions are validated server-side against an allowed-transition set, so a stale
device cannot drive an inspection backwards.
7.3 What we deliberately do not build
No merge UI for text, no attempt to auto-merge prose, no block-level structural merge, no
"apply the rest for me" bulk action. Product ruled out a collaborative editor and the
conflict surface is kept correspondingly small and honest.
8. Attachments (photos, signature images)
8.1 Capture
Local photo_id (UUIDv7) generated at shutter time; it is the idempotency key, so a retried
capture never duplicates.
SHA-256 of the content for integrity and dedupe.
EXIF orientation normalised at capture; capture timestamp preserved from the device clock
alongside a monotonic device sequence, so clock skew cannot reorder history.
No location metadata — these are indoor inspections, so geotagging is needless
personal-data collection.
Quota check before capture: the 40-photo per-inspection cap and a device free-space
threshold. Capture is refused with a clear message rather than failing silently later.
8.2 Upload protocol
Separate binary channel. Photo bytes never enter the JSON op payload.
POST /v2/inspections/{id}/photos with photo_id → idempotent create → upload_id and
chunk size.
PUT /v2/uploads/{upload_id} with Content-Range and a per-chunk checksum.
GET /v2/uploads/{upload_id} → received bytes, for resume.
POST /v2/uploads/{upload_id}/commit with the whole-file SHA-256 → server verifies →
returns photo_ref and the new record version.
Per-photo state machine: local → uploading(partial, offset) → uploaded → committed. Resume
is from the byte offset, so a kill mid-upload resumes rather than restarts. Committed-but-not-
yet-referenced uploads are held server-side as pending and garbage-collected after a documented
window, so a device that dies before its submit op does not leak storage.
8.3 Retention and purge (the 24-hour rule)
Deadline is computed as server_accepted_at + 24h, using the server timestamp (A7). The
device never starts its own clock at acceptance; it starts it at the moment it learns of
acceptance, but measures against the server's time so a late-delivered notification does
not silently extend the window.
Two-tier retention after purge:
Purged — photo files, signature image and strokes, free-text answers. This is
customer data under the assumed reading (A8).
Retained — a minimal non-customer receipt: inspection id, photo ids and hashes,
signature hash and timestamp, submit/accept timestamps. This supports the audit history
and compliance evidence without holding customer content.
The device cannot rely on being online to receive the acceptance. If the local clock
passes a known deadline, the device purges on its own initiative at the next opportunity —
app launch, foreground, or scheduled work — and logs the purge as a compliance event.
If the device is offline past the deadline and the acceptance state is still unknown, the
safe default is applied to anything already submitted (the server holds that copy).
Unsubmitted work is a separate question and is explicitly an open decision (D3) — it is the
device's only copy, and destroying it on an unverified assumption would be worse than the
alternative. This must be confirmed by legal, not decided here.
MDM remote wipe is the backstop for a device that never returns to coverage; it is a
complement to, not a replacement for, the local purge.
9. Security
Device posture (A4): company-managed tablet, known MDM, minimum OS version enforced,
app container excluded from OS backup, screen lock required, remote wipe available. Root or
jailbreak detection with a defined consequence — most defensibly, blocking offline mode
rather than blocking the whole app, so a flagged device can still read and submit.
Identity and transport: short-lived access token with a rotating refresh token in the
platform keystore; long-lived credentials never in SQLite. TLS 1.2+ only, certificate pinning
to the company CA, and mutual-TLS device certificate for device identity and attestation.
At rest: SQLCipher for the database, with the data-encryption key held in the hardware
keystore. Photo and signature files encrypted with a key derived per user and per device, with
a re-key path on enrolment change. Key scope is open decision D4 — per-device is simpler,
per-user is correct if tablets are shared.
Authorisation: role checks (inspector vs supervisor) and per-inspection ownership are
enforced server-side on every sync endpoint. The device is untrusted. Explicit IDOR
testing in §10.
Audit integrity: audit entries are append-only and server-signed. A device cannot author a
trusted entry, only a provisional one (§6.5).
Privacy: no location, contacts, or advertising SDKs. Analytics and logs carry no field
values, no photo bytes, and no free text (§11.3). The signature screen masks the app-switcher
snapshot and blocks clipboard copy. Signature binding to a content hash means post-hoc content
edits are detectable rather than invisible.
Offboarding: unenrolment destroys keys and triggers remote wipe; server-side retention
policy is separate from the device's 24-hour rule and is out of this plan's scope.
10. Migration and backfill
The good news, if A10 holds: the existing store is a read-through cache, so it should hold
no durable user-authored content and there is little to backfill. This must be verified,
not assumed — a partially shipped offline draft feature or a user relying on cached text would
turn this into a real backfill workstream.
Schema versioning: PRAGMA user_version plus an applied-migrations table. Forward-only,
idempotent steps, each verified in CI. No automatic downgrade; a device that cannot migrate
forward takes the defined rebuild path.
Rebuild path: on unrecoverable migration failure, the app offers a logged rebuild. It
refuses to wipe while the outbox is non-empty unless the user makes an explicit, recorded
decision — a silent wipe of in-flight field work is the failure mode this project is meant to
eliminate.
Backfill of assignments: pull the assignment index in full, then details lazily per
inspection, so first sync on a large assignment list stays fast on a weak link.
Cutover: an offline_writes flag scoped by organisation/cohort. With the flag off, the
sync engine is idle and the old read-through path is unchanged — the read path remains
testable until the old path is deliberately removed.
Existing devices on the old app version must keep working against the new server. The
server is backward compatible with at least one previous app version for the whole pilot.
Verification: after migration, run the §5.4 invariant check on a seeded dataset and
compare read results against the old cache to prove the read path did not change.
11. Observability
11.1 On-device
Counters and timings, written to local structured storage and flushed as metadata only:
outbox depth and oldest pending op age (the one that predicts "an inspector lost work")
photos pending upload, total pending bytes
conflicts detected / resolved, and by class
purge obligations due, overdue, and completed
last successful sync time, feed token age
database and file store size, free space
crashes during a sync cycle, quarantined (poison) ops
Surfaced in-app for the inspector's own device health ("12 items waiting to upload, oldest 3h")
and readable by support.
11.2 Server
per-op apply latency and error class, conflict and rejected rates
delta feed size, round-trip count, token-expiry/full-resync frequency
photo upload throughput, resume frequency, hash-mismatch and dedupe rates
purge acknowledgements; overdue purge alerts as a compliance page
devices that have not connected within a threshold window
11.3 Rules
No customer content in telemetry — no field values, no free text, no photo bytes, no
signature images. Correlation by op_id and inspection_id only. Enforced by an automated
redaction scan in test (§10).
Alerts: oldest pending op age above threshold; purge deadline breach; photo upload failure
rate; conflict-rate spike (a spike means a real workflow problem, and is the signal that
A9 or the §7.1 policy is wrong).
Support tooling: a server-side view of a device's last-seen version and pending count, and a
re-push of an inspection to a device — so support can diagnose a stuck device without
physical access to it.
Explicit non-goal: no per-keystroke or per-field analytics.
12. Milestones
Sequenced by technical dependency, not by calendar. Two tracks (device and server) can run in
parallel once M1 is agreed, but each track needs its own accountable owner; this plan does not
assume a particular team size or composition.
ID
Milestone
Depends on
Acceptance criteria
M0
Decisions closed
Product, compliance, security input
Every decision in §13 marked decided with an owner and a recorded outcome. Compliance has explicitly accepted or rejected the §8.3 purge interpretation. A14/version-enforcement behaviour is confirmed in writing by the server owner.
M1
Sync + merge contract spec
M0
Written spec covers the op envelope, every response class, the delta feed, token expiry, version semantics, and the merge rule per data class. Reviewed and signed off by both track owners. A shared contract-test suite exists and fails against a stub server.
M2
Durable local store + transactional outbox
M1
With the server unreachable and networking fully disabled, a full inspection — 40 photos, signature, answers, submit — can be created, edited and submitted, and survives app kill and device reboot. Automated crash injection at every write step shows no loss. Read path shows no regression.
M3
Server sync endpoints (flagged off)
M1
Contract tests green. Replaying the same op_id yields identical server state (idempotency). A stale-version mutation is rejected with the current record returned. No production traffic reaches the endpoints while the flag is off.
M4
Push/pull engine, merge, conflict records
M2, M3
Merge property tests hold the invariants: no scalar field is lost without a conflict record, merge is deterministic, re-applying an op is idempotent, and two divergent replicas converge. Two-replica divergence test passes for every data class.
M5
Attachment channel
M2, M3
40 photos upload and resume correctly from an arbitrary byte offset, including after a mid-upload kill, on a constrained and flaky link, within the agreed time budget. Server rejects truncated, corrupted, and hash-mismatched blobs. Duplicate photo_id does not duplicate.
M6
Compliance purge and retention
M0, M2, M4
Purge completes within the deadline measured from the server acceptance timestamp, in every test variant including the device-offline-past-deadline case. Purge produces complete, exportable evidence. Zero customer payloads remain past the deadline in an automated audit. Legal has ruled on D3.
M7
Security hardening
M1
Pen-test findings triaged to zero high/critical. Enrolment, unenrolment, remote wipe and key rotation all exercised in staging. Jailbroken-device behaviour verified. Log redaction scan passes against seeded field values.
M8
Observability and support tooling
M1, M3
Dashboards live for the §11 metrics. Alerts fire in a rehearsed game-day. A support engineer can diagnose a stuck device and trigger a re-push without device access. Telemetry confirmed content-free.
M9
Field pilot enablement and UAT
M4, M5, M6, M7, M8
The scripted offline UAT passes on target tablet hardware with a real inspector: complete a 40-photo signed inspection fully offline, kill the app twice, reconnect, and the server record matches. Runbook, escalation path, and kill switch documented and rehearsed.
M10
Pilot readout and rollout decision
M9
Gate metrics below measured and reported per cohort. An explicit go / hold / rollback decision recorded with rationale. Per-cohort flag configuration exists for expansion, and the old path is still intact if rollback is chosen.
12.1 Rollout stages and gates
Stages are gated on measured thresholds, not on elapsed time.
Stage
Configuration
Exit gate
0 — Internal
Dev builds, synthetic network shaping, real basement hardware
M2/M3/M4 acceptance green
1 — Shadow
Sync engine runs; app still writes through the old path; engine output is compared, read-only, against what the old path produces
Zero unexplained divergence between pulled data and the old read path
2 — Pilot
Flag on for a small consenting cohort (one supervisor team)
All of: zero confirmed data-loss reports; ≥ the agreed share of pilot inspections submitted via the offline-first path; p95 oldest-pending-op age under threshold; purge compliance 100%; conflict rate below threshold with every conflict class observed and understood
3 — Expand
Enable per team/region, same gates each time
Each expansion independently passes the Stage 2 gate
4 — Default
Flag default-on; old path removed only after Stage 3 is clean
No regression in the funnel metrics; old path removal is a separate, reversible change
5 — Review
Post-pilot decision on deferred items
Deferred-items table (§4) revisited with real data
"Zero confirmed data-loss reports" is the gate that cannot be traded. A threshold miss elsewhere
is a hold; a single data-loss report is a rollback.
12.2 Rollback
Flag off returns devices to the read-through path. The outbox is not discarded: it is
parked (state = parked) and remains readable.
If the flag goes off with a non-empty outbox, the app must either drain it first or park it
explicitly and visibly — never silently drop it.
The server must accept ops from a parked cohort whenever the flag returns on, so rollback is
genuinely reversible. This is a hard server-side requirement, not a nice-to-have, and it is
an acceptance criterion of M3.
Kill switch: a server-side ability to reject ops from a specific app version with a
graceful, user-readable error, for a data-corrupting bug that must be stopped immediately.
Schema migration rollback: never auto-downgrade. Migrations are idempotent and
forward-fixable; a bad migration is fixed forward.
Rollback is rehearsed during M9, not documented and hoped for.
13. Decisions required before implementation
Each has a recommendation. Recommendations are my judgement given the brief, and the ones
marked (blocking) stop M1 if unresolved.
#
Decision
Recommendation
Why / what it costs
D1(blocking)
How is a same-field concurrent edit resolved?
Server value canonical, device value preserved as a conflict record, supervisor resolves in the web app. No blocking for the inspector.
Keeps the audit requirement and keeps inspector work. Costs a resolution queue — the pilot must measure its size. See §7.1.
D2(blocking)
Can the server be extended additively, and is the integer version actually enforced (A2/A14)?
Yes — additive endpoints plus If-Match-style conditional mutation on the existing integer version.
If the server cannot change, this becomes a separate sync service and the scope of M1–M4 changes materially. If the version is not enforced, three-way merge has no trustworthy base.
D3(blocking, legal)
Do unsubmitted local drafts count as customer data under the 24-hour rule?
Exclude never-submitted drafts from the 24-hour clock — they are the only copy and not yet part of the inspection record — but cap their retention and make them explicitly user-deletable.
Purge-first destroys the inspector's only copy. Purge-never risks a violation. This is a compliance determination, not an engineering one.
D4
Local encryption key scope: per-device or per-user?
Per-user keys, server-wrapped, with re-key on enrolment change.
If each tablet is permanently assigned to one inspector, per-device is simpler and acceptable. The answer depends on the device-sharing policy, which the brief does not state.
D5
Does a submitted-but-unaccepted inspection block acceptance until its photos finish uploading?
No. Allow submit while photos are pending, with a clear outstanding count; let acceptance wait server-side until required content is present.
Otherwise an inspector in a basement cannot leave, because 40 photos cannot upload without signal. This is the difference between a usable feature and a trap.
D6
Photo custody path: is there existing object storage, and do supervisors need photos on the web immediately?
Reuse existing attachment storage behind an additive upload-session endpoint. Do not build new blob infrastructure.
Building new storage is a large, unrequested workstream.
D7
Signature format: image + metadata, or PKI/certificate signing?
Captured stroke image + audit-bound content hash + signatory attestation, consistent with current e-signature practice. No PKI.
Certificates add PKI operations, revocation handling, and trust infrastructure, and give nothing the brief asks for.
D8
Do supervisors get blocked while a conflict is open?
No. They see it and can accept the authoritative value or pull the device variant.
Blocking the supervisor record would let one offline tablet stall a review queue.
D9
What happens to a photo the inspector deletes after upload?
A tombstone op; the deletion is an auditable event.
Erasing the server object would break the audit history the regulations require. Needs compliance agreement.
D10
Does the device need offline access to historical/accepted inspections, or only assigned open ones?
Only assigned open inspections, plus a minimal non-customer receipt.
Offline history expands the local store and the purge surface dramatically. Reopen only on a concrete supervisor request.
D11
Form schema changes while a device is offline.
Version the form definition; preserve answers to removed fields in an orphaned bucket and surface them; never silently discard a filled field.
Dropping answers an inspector typed is a data-loss event even if the field no longer exists.
D12
How long may the device hold unsubmitted work?
No hard cap, but a staleness rule: an assignment untouched for an agreed period prompts "is this still current?" — no auto-expiry.
Auto-expiry destroys field work; no prompt at all leaves inspectors submitting against stale assignments.
D13
Is at-rest encryption (SQLCipher + keystore) approved, and is FIPS required?
Confirm early; it is a long-lead item. If FIPS is required, the crypto choice changes and affects both the database and file storage.
Late discovery of a FIPS requirement invalidates the storage design in M2.
D14
Target platform and its background-execution limits (A13).
Design for foreground-only sync; treat reliable background execution as an optimisation, never a dependency.
Depends on the platform, which the brief does not name.
D15
Conflict visibility on the inspector's device.
A non-blocking "needs supervisor review" badge with the preserved value; never block the next inspection.
A blocking dialog strands a working inspector for a problem they cannot fix.
14. Testing
14.1 Offline and durability
Network shaping: full offline, 200 ms–5 s RTT, packet loss, bandwidth caps, and signal
flapping mid-upload.
Crash injection: kill the process at every step of the write path (M2's core evidence) and
at every upload chunk boundary; verify no answer, photo, or signature is lost.
Reboot between capture and sync; low storage; battery saver; clock skew on the device.
Multi-day darkness (A6): assignment still opens, staleness prompt appears, no data loss.
14.2 Merge and protocol
Property/fuzz tests on the merge engine over randomly generated two-replica edit
sequences, asserting the invariants: no scalar field lost without a conflict record; merge is
deterministic; re-applying an op is idempotent; divergent replicas converge; delete is never
resurrected.
Idempotency: same op_id applied twice yields one effect and the same response.
Version conflict returns the current record; a stale device cannot move status backwards.
Delta feed correctness under concurrent server-side writes; token expiry forces the defined
full-resync path and it completes cleanly.
Cross-device cases (A5): two tablets, one inspection, both offline, then reconnect.
Contract tests shared between client and server so schema drift fails CI rather than
production.
The §5.4 invariant is a test, not a comment.
14.3 Attachments
Chunked upload resume from every offset, including after a mid-upload kill.
Truncated, corrupted, and hash-mismatched blobs are rejected; duplicate photo_id does not
duplicate; out-of-order or duplicated chunks are handled; storage-full mid-upload recovers.
The old read-through path stays under test until Stage 4, so its removal is a deliberate
decision rather than an untested assumption.
No real customer data in any test environment — synthetic photos, synthetic names.
14.7 Field validation
Scripted UAT (M9) on real tablet hardware, in a real basement, with a real inspector and a
timed task.
The acceptance scenario in one line: complete a 40-photo signed inspection fully offline,
kill the app twice, reconnect, and the server record matches what was captured.
Performance budgets set before the pilot: first assignment sync time on a weak link, delta
round-trip count, and 40-photo upload time.
15. What this plan deliberately does not promise
No dates or durations. Sequencing is dependency-driven; scheduling needs the team's actual
composition, which the brief does not provide.
No headcount. Each milestone names an accountable role, not a number of people.
No claim that the conflict queue is free. §7.1 is the right trade for a no-data-loss
requirement, and it creates real operational work that the pilot must quantify.
No assumption that connectivity is good, that devices are online, or that the server
behaves better than A2 describes. Each of those is a decision or a gate, not a premise.
mcode-m3.1-flash/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.5 / 10graded blind as submission C
A concise plan that covers everything, with sensible defaults: autosave as the source of truth, field auto-merge plus inspector choice, a supervisor-authoritative denylist, cryptographic erasure, and 'accepted with pending changes'. The 24-hour purge is keyed to when the tablet observes acceptance and treats late observation as a compliance exception, so the offline-device gap is acknowledged but not closed.
Strengths
Clear outbox with leases, a submit barrier, and resumable chunked uploads
Per-inspection keys allow cryptographic erasure
Eight decisions, each with a recommendation and reasoning; D4 handles accept-while-pending
The rollback data-safety invariant is stated explicitly
Weaknesses
Purge starts at observation; a device offline past the window is only logged as an exception
No dedicated server-foundation milestone (it is folded into M3 dependencies), and observability arrives late in M6
Pilot thresholds are deferred to D8
Evidence the grader checked
§9 24-hour purge: 'If the tablet stays offline past the window, purge on next launch/sync and record the delay as a compliance exception'
§15 D7 recommendation
§14 M1–M8 sequence
Objective checks
Files
muse/06-product-planning/RESPONSE.md
Riskiest assumptions — offline inspection sync
The PLAN.md assumptions with the highest combination of uncertainty and
blast radius are A2 (edit overlap), A4/A5 (photo upload under intermittent
connectivity), and A9 (24-hour purge). Each is summarized below with why it
is risky, what breaks if it is wrong, and how the plan hedges.
1. Supervisors rarely edit the same form fields an offline inspector is editing (A2)
Why it is riskiest. The entire conflict policy — auto-merge
non-overlapping fields, guided per-field choice for true conflicts, no
collaborative editor — rests on overlap being uncommon. The brief only says
supervisors "can edit the same inspection from the web"; it says nothing
about which fields or how often. If supervisors routinely correct answers
while inspectors are offline in basements, conflicts become the common case:
inspectors face a conflict screen after every shift, submissions stall, and
either side's work is at risk of being bulk-overwritten under time pressure.
If wrong. Conflict-resolution UX becomes the product, not the edge case;
"no silent overwrites" still holds, but throughput and trust collapse, and
supervisor accept/reject queues clog on "pending tablet changes."
Hedge in the plan. Per-field diff with both values/authors/timestamps,
bulk + per-field choices, supervisor-authoritative denylist for status
fields, losing-side preservation in audit, web warnings before editing
records with pending tablet drafts, and a conflict-rate dashboard with pilot
exit gates (decisions D2, D4, D8). Pilot must measure true-conflict rate
before broad rollout; if it is high, the revisit is workflow (who edits what,
when), not a bigger merge algorithm — product already ruled out a full
collaborative editor.
2. Forty photos can be uploaded over intermittent connectivity within an acceptable window (A4 + A5)
Why it is risky. Photos dominate bytes, time, and failure modes. Forty
multi-MB originals over basement-grade links with mid-batch interruptions
may never converge: uploads restart without resume, storage fills, submits
block on missing attachments, and inspectors learn to delay or skip photos.
The brief fixes the count (40) but not sizes, quality obligations, or how
long a usable window lasts — all three multiply into one throughput gamble.
If wrong. Inspections pile up in "photos pending," submit barriers never
clear, tablets run out of encrypted storage, and retry storms on brief
reconnects crowd out field-op sync.
Hedge in the plan. Resumable chunked uploads with per-chunk ack and hash
verification, independent per-attachment retry, prioritized signature/field
sync, local 40-photo enforcement with visible counters, low-storage stress
tests — and, decisively, decision D3: define a compressed rendition as the
record if compliance permits. The pilot gate on attachment completion rate
is the falsification test for this assumption.
3. "Removal within 24 hours after acceptance" is achievable from a tablet that may be offline (A9)
Why it is risky. The obligation is regulatory, but the tablet cannot act
on an acceptance it has not observed. If "24 hours after acceptance" is read
as server-side wall clock with no tolerance for offline delay, every
tablet offline for a day after acceptance is a breach by construction — and
thumbnails, caches, and outbox payloads widen the purge surface beyond the
obvious files. The brief mandates the outcome but not the clock, the scope
of "local customer data," or the accounting for offline delay.
If wrong. Silent non-compliance: customer data lingers past the window
with no record, or conversely an over-aggressive purge deletes unacked
drafts/photos to meet a misread deadline.
Hedge in the plan. Purge scheduler keyed to observation with reboot- and
background-safe execution, cryptographic erasure, purge scope covering
answers/photos/signature/thumbnails/caches/outbox, purge_log receipts, and
— critically — decision D7: server acceptance starts the obligation, the
tablet purges within 24 hours of observation, and late observation is logged
as an auditable compliance exception rather than silent breach. Legal and
compliance must confirm D7 before implementation; telemetry alerts on any
purge-window breach during pilot.
muse/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
muse/06-product-planning/PLAN.md
Offline inspection sync — Plan
1. Goal and context
Enable inspectors on company-managed tablets to complete assigned building
inspections fully offline — open assignment, fill forms, capture up to 40
photos, collect a signature — and sync reliably when connectivity returns
(intermittent or delayed), without losing notes.
Current state (per BRIEF.md): losing connection after opening an inspection
can discard notes. The mobile app has SQLite but uses it only as a
read-through cache. Server APIs are REST with integer record versions.
Supervisors can edit the same inspection from the web. Regulations require an
audit history and removal of local customer data within 24 hours after an
inspection is accepted. Product wants a pilot before broad rollout and
explicitly does not want a full collaborative editor.
This plan is the implementation blueprint. It makes no staffing or calendar
commitments.
2. Explicit assumptions
Assumptions are numbered so milestones, tests, and decisions can reference
them. Items marked must-confirm are in Section 13 as pre-implementation
decisions.
A1. Assignment model. Inspections are assigned to one inspector at a
time; reassignment is rare. Offline work is single-writer per inspection on
the tablet side. Concurrent offline edits from two tablets on the same
inspection are out of scope.
A2. Overlap with supervisor edits. Supervisors edit from the web mainly
status/metadata (assign, accept, reject, priority, due date, notes to
inspector), and only occasionally inspection form answers. Field-level
overlap during an active offline session is uncommon. Must-confirm.
A3. No real-time collaboration. Product accepts that offline inspectors
see a snapshot from last sync; they do not need live supervisor edits while
offline, and supervisors do not need live inspector keystrokes. Sync
converges on reconnect.
A4. Connectivity profile. Offline periods last minutes to a full shift
(basements, intermittent return). The protocol must assume sync can be
interrupted mid-batch and resumed later; it must not assume a long stable
window. Must-confirm (affects attachment strategy).
A5. Photo volume and size. "Up to 40 photos" per inspection; tablet
cameras produce multi-MB images. Originals must be preserved or a defined
compressed equivalent accepted as the record. Must-confirm.
A6. Signature. One signature per inspection, captured on-device,
treated as a small binary attachment bound to a form version, not as
editable text.
A7. Server versioning. REST integer record versions are per-inspection
(or per-record) optimistic-concurrency counters: a write with a stale
version is rejected (e.g., 409) and the client must re-read and retry. The
server can add idempotency keys and per-field audit metadata without a
full rewrite. Must-confirm.
A8. Audit history. "Audit history" means a server-side append-only log
of who changed what field, when, from which device/client, sufficient for
regulators; the tablet does not need to render full history offline, only
to contribute accurate entries. Must-confirm.
A9. 24-hour local deletion. "Removal of local customer data within 24
hours after an inspection is accepted" applies to form content, photos,
signature, and derived caches/thumbnails on the tablet, measured from the
tablet learning of acceptance (or from acceptance server-side plus a bounded
propagation window — decision required). MDM remote wipe exists as a
backstop but is not the primary mechanism. Must-confirm.
A10. Device trust. Company-managed tablets provide OS-level storage
encryption, screen lock, and MDM inventory/wipe. The app adds its own
encrypted store and does not rely on OS encryption alone for PII.
A11. Auth offline. Inspectors authenticate while online; short offline
grace (cached credentials / tokens) is acceptable to regulators and
security, with re-authentication required for sensitive actions per policy.
Must-confirm.
A12. SQLite stays. The existing SQLite database can be migrated to a
durable offline store (schema + outbox tables); no change of local database
engine is needed.
A13. Pilot scope. Pilot is a small set of inspectors/building types with
cooperative supervisors, on the same server version as production (behind
feature flags), not a fork.
3. User flow
3.1 Inspector (tablet)
Before going offline (online). Inspector sees assigned inspections.
Opening (or explicitly downloading) an assignment pins it for offline:
form schema, current answers, reference data, and record version are
persisted. UI shows "Available offline" with last-synced time.
Going offline. Banner changes to "Offline — changes saved on this
tablet." All input autosaves locally on every change (no save button as
the only path; today's data-loss bug is fixed by making local persistence
the source of truth for drafts).
Offline work. Fill forms (including required-field validation run
locally), take up to 40 photos (queued with thumbnails), collect
signature. Draft state, photo queue, and completion checklist are visible
without connectivity. App restart, OS kill, or reboot loses nothing.
Reconnect (intermittent or full). Sync starts automatically and in the
background; UI shows per-inspection status: Not synced / Syncing (n/m photos) / Synced / Conflict needs review / Failed — will retry. Inspector
can keep editing during sync; new edits queue behind the in-flight batch.
Conflict. If a supervisor changed the same record while the inspector
was offline, the inspector is shown a conflict screen: which fields
differ, supervisor value vs. tablet value, and a guided choice (keep mine /
take supervisor / merge per field where safe — see Section 6). No silent
overwrite in either direction. Unresolved conflicts block "submit" but
never block continued local editing.
Submit / accept. Inspector submits when complete (works offline; the
submission is queued). Supervisor accepts on the web. Once the tablet
learns of acceptance, a deletion countdown starts and local customer data
for that inspection is purged within the 24-hour obligation (Section 8),
leaving only a non-customer receipt (inspection ID, timestamps, sync
status) if permitted.
3.2 Supervisor (web)
Web shows sync/liveness per inspection: tablet draft pending,
last tablet sync, photos pending (n/m), conflict, so supervisors do
not mistake a stale view for final.
Supervisors can keep editing; if they edit a record with a pending tablet
draft, the UI warns "inspector has unsynced changes; your edit may
conflict" (best-effort, based on last known tablet state).
Accept/reject remains a server-side transition; rejection returns the
inspection to the inspector's queue with a reason, deliverable on next sync.
4. Scope boundaries
In scope:
Offline open, edit, autosave, and resume for assigned inspections.
Photo capture queue (40 per inspection), signature capture, background
sync with resume.
Conflict detection and guided resolution for the tablet-vs-web case.
Audit events for offline-originated changes; 24-hour local purge after
acceptance.
Pilot instrumentation, feature flags, staged rollout, and rollback.
Out of scope / non-goals:
Full collaborative/real-time co-editing (explicitly excluded by product).
Multi-tablet concurrent offline editing of the same inspection.
Offline supervisor workflows on tablet; offline reassignment; offline
creation of new unscheduled inspections (unless D-decision says otherwise —
default: not in pilot).
General offline browsing of unassigned/history inspections beyond pinned
assignments.
New server framework or protocol replacement (e.g., WebSockets, CRDT
service); the design extends REST + integer versions.
Pilot exclusions (add after pilot only if validated): bulk pre-download
of a week of assignments, offline reference-data search across all buildings,
background photo upload on metered connections beyond a cap.
5. Local data model (SQLite + file store)
Keep SQLite; promote it from read-through cache to durable store. All writes
go to SQLite first, then to an outbox.
Proposed tables (names illustrative; final schema in implementation design):
inspections (one row per pinned assignment): id (server ID),
assigned_to, status (local lifecycle: pinned/draft/submitted/ conflict/accepted-purge-pending/purged), form_schema_version,
server_version (last synced integer version), local_revision
(monotonic local counter), updated_at_local, last_synced_at,
last_server_state_json (snapshot for diff/conflict display),
acceptance_observed_at.
form_answers: (inspection_id, field_key) with value_json,
updated_at_local, updated_by, base_server_version.
outbox_ops (append-only queue): op_id (client UUID),
inspection_id, type (upsert-field / submit / resolve-conflict),
payload_json, base_server_version, created_at, attempt_count,
state (pending/in-flight/acked/failed), idempotency_key.
In-flight ops carry a lease so a crash returns them to pending.
attachments: attachment_id (client UUID), inspection_id, kind
(photo/signature), local_path, byte_size, hash_sha256,
capture_at, upload_state (queued/uploading/parts-acked/complete/
failed), server_blob_id, form_version_binding (for signature).
Binaries live in an app-private encrypted file store; SQLite holds only
metadata.
Direction split. Pull-then-push per inspection, per cycle:
Pull: GET /inspections/{id}?since={last_pull_version} returns current
server version, changed fields, supervisor transitions, and upload
session state for attachments.
Push: POST outbox ops in creation order, each with base_server_version
and idempotency_key (client op UUID). Server applies if versions
chain, else rejects with 409 + current server state.
Idempotency. Server stores idempotency keys per inspection; retried
POSTs return the original result without duplicating effects or audit
entries. Required for at-least-once mobile retry.
Batching and resume. Push in small batches (e.g., one field-batch +
one attachment chunk-set per request; exact sizes tuned in pilot). Each
batch is independently acked; on interruption the client resumes from the
first unacked op. No "all or nothing" multi-photo transaction.
Ordering. Field ops apply in client creation order per inspection.
Attachment uploads are independent of field ops and can interleave, but
submit is a barrier op: it applies only after all preceding field ops
and required attachments for that inspection are acked.
Submit semantics.submit carries the local revision and required
checklist hash; server validates completeness against the authoritative
schema version and either accepts (status transition + version bump) or
rejects with machine-readable errors the tablet can render offline later.
Backoff and triggers. Exponential backoff with jitter on failure;
sync triggers on connectivity regain, app foreground, edit, periodic
while online, and manual "Sync now." Never busy-loop on metered or
flaky links; respect OS background limits.
Schema versioning. Form schema is versioned; pinned inspections record
the schema version they were opened with. If the server schema advanced
during offline work, push includes the base schema version and the server
either migrates forward (additive changes) or returns a
migration-required response with instructions (breaking changes).
7. Conflict policy (no collaborative editor)
Detection. Any push whose base_server_version no longer matches the
server version is a conflict candidate. The server responds 409 with the
current field values and their versions/authors; the client diffs against
last_server_state_json to classify per-field: clean (only one side
changed), true conflict (both changed to different values), or
metadata-only.
Default resolution rules (proposed; confirm in D-decisions):
Same-value writes to the same field auto-resolve (no user action).
True field conflicts require explicit choice on the tablet (inspector
keeps mine or takes supervisor), except a small denylist of
supervisor-authoritative fields (e.g., accept/reject, assignment) that
always resolve to the server value and are shown as read-only during
conflict.
Photos and signature are append-only / replace-with-history: concurrent
additions union; a replaced signature creates a new version rather than
overwriting, with the old version retained server-side in audit.
Never lose data. Losing side of a conflict is preserved in server audit
and (until purge) in the tablet conflict snapshot, restorable by
supervisor action.
UX. One conflict screen per inspection listing each conflicting field
with both values, author, and timestamp; "accept all supervisor" /
"keep all mine" bulk actions plus per-field pickers; resolution itself is
an outbox op (works offline after the conflict snapshot is known, applies
on next push).
Server guardrails. Supervisor web edits to an inspection with known
pending tablet changes show a warning; accept/reject while tablet ops are
unacked is either blocked or recorded as "accepted with pending tablet
changes" per decision D4.
8. Attachments (photos + signature)
Capture. Up to 40 photos per inspection (enforced locally with a clear
counter). Each photo gets a client UUID, SHA-256, timestamp, and
inspection binding at capture. Signature is one attachment bound to the
form revision signed; editing signed fields after signing invalidates the
signature locally and requires re-sign (configurable per compliance).
Storage. App-private encrypted file store; originals retained until
server ack (or per compression decision). Thumbnails derived locally for
list views; thumbnails are part of the purge scope.
Upload. Resumable chunked upload (e.g., client-initiated upload
session per attachment, fixed-size chunks with per-chunk ack, whole-file
hash verified server-side before marking complete). Retries resume at
chunk granularity. Photo order does not matter; signature upload is
prioritized after field ops.
Compression/quality (decision D3). Either upload originals, or
upload a defined compressed rendition as the record with originals
discarded after ack. The choice affects storage, upload success on flaky
links, and regulatory acceptance — must be decided with compliance before
implementation.
Failure handling. Failed attachments retry independently; an
inspection can reach Submitted with attachments pending only if policy
allows partial submit (default recommendation: submit requires all
required attachments acked; optional photos may follow within a grace
window — confirm).
9. Security
Encryption at rest. SQLite + file store encrypted with keys in the
platform keystore/keychain, unlocked on device unlock; per-inspection or
per-file data keys so purge can be cryptographic erasure (destroy key +
delete bytes).
Encryption in transit. TLS; certificate pinning per company policy.
Auth. Online login via existing IdP; short-lived access tokens with a
bounded offline grace period (decision D6). Jailbreak/root detection and
MDM compliance checks gate offline data access where policy requires.
24-hour purge. On observing acceptance (push ack, pull, or push
notification while online), record acceptance_observed_at, schedule
purge, and delete form answers, photos, signature, thumbnails, outbox
payloads, and audit staging for that inspection within the remaining
window — including when the app is backgrounded (OS background task) and
after reboot (purge check on launch). Retain only purge_log receipt.
If the tablet stays offline past the window, purge on next launch/sync and
record the delay as a compliance exception with telemetry.
Backstops. MDM selective wipe for lost devices; supervisor-initiated
remote revoke that queues a purge for next contact.
Audit. Every mutation carries actor, device ID, client timestamp,
base version, and op ID; server appends to the immutable audit log.
Client clocks are untrusted for ordering — server timestamps are
authoritative; client timestamps are informational.
10. Migration and backfill
Client migration. Existing read-through-cache SQLite migrates in place:
add outbox/attachment/sync/purge tables, backfill server_version from
last cache metadata, and treat any unpersisted in-memory drafts as lost
(one-time, disclosed in release notes — the new autosave prevents
recurrence). Migration runs in a transaction with rollback to the old
schema version on failure.
Server migration. Additive changes only for pilot: idempotency-key
store, per-field audit metadata, resumable upload sessions, conflict
(409 + state) responses, pending tablet changes signals for web UI.
No change to version semantics; no data rewrite.
Backfill. No historical backfill of audit beyond existing records.
In-flight inspections at rollout cutover sync normally on next contact;
inspectors are asked to sync before updating the app (banner + release
notes), but update without sync must not corrupt data (old cache rows are
ignored, never pushed).
Downgrade. App downgrade across the migration is unsupported during
pilot; forward-rollback (Section 12) is the path.
11. Observability
Device telemetry (privacy-preserving). Sync success/failure rates,
conflict rate and resolution outcome, outbox depth and age, photo
queue depth and chunk-retry counts, time-to-synced after reconnect,
purge latency vs. 24-hour window, storage pressure, migration success.
No customer PII in telemetry; IDs are inspection/op UUIDs only.
Inspector-facing. Per-inspection sync status, last-synced time,
pending counts, conflict entry point, purge countdown after acceptance.
Supervisor-facing. Pending-tablet-changes and conflict badges,
last-tablet-sync timestamps.
Pilot dashboards and alerts. Daily pilot review of sync success,
conflict rate, attachment completion, purge compliance, and crash-free
sessions; alerts on purge-window breaches, audit gaps, and sync stalls
(outbox age beyond threshold).
12. Staged rollout and rollback
Stage 0 — Internal validation. Team devices on lab networks with
scripted offline/chaos cases; all acceptance criteria in Section 14
exercised; no customer data beyond fixtures.
Stage 1 — Pilot. Small inspector cohort + paired supervisors, feature
flags per device/inspection type; server changes behind flags with old
behavior default. Entry criterion: Stage 0 acceptance green. Exit
criterion: pilot success thresholds (Section 14) met over a sustained
window plus supervisor sign-off.
Stage 2 — Broad rollout. Progressive enablement by team/region with
pause criteria (purge breach, spike in 409s/unresolved conflicts, upload
completion drop, crash regression). Each step is reversible via flags.
Rollback.
Server: flags revert to pre-offline behavior; in-flight upload sessions
drain or expire; unacked tablet outbox ops remain queued and retry
against the reverted API only if the API contract is unchanged —
otherwise they hold with clear UI until re-enabled (never silently
dropped).
Client: forward-fix preferred; if a client build must be pulled, MDM
holds the rollout and affected tablets keep local drafts intact (no
wipe on rollback). Downgrade across the SQLite migration is not
supported; recovery is via fixed build, not data deletion.
Data safety invariant: rollback never deletes unacked tablet drafts or
unacked photos; purge obligations still apply to accepted inspections.
13. Testing
Unit. Outbox state machine (pending/in-flight/acked/failed + lease
expiry), conflict classifier (clean/auto-merge/true-conflict), purge
scheduler math, chunk resume logic, idempotency-key handling.
Integration (client + API contract). Pull-push cycles against a fake
server implementing version/integer + 409 semantics; replayed idempotency;
submit barrier with missing attachments; schema-version migration paths.
Persistence/crash. Kill -9 / reboot at every point (mid-edit,
mid-photo, mid-chunk, mid-batch, mid-purge); assert zero draft loss and
correct resume. Downgrade-block and migration-failure rollback cases.
Network chaos. Scripted flaky/intermittent profiles (drop mid-batch,
reorder, duplicate POSTs, long offline then burst reconnect); duplicate
delivery must not duplicate effects or audit entries.
Conflict matrix. Supervisor-edits-field × inspector-edits-same-field,
supervisor-accepts-while-tablet-pending, rejection-with-pending-draft,
signature-invalidated-by-post-sign-edit, schema-bump-mid-offline — each
with expected UI and server end-state.
Attachment stress. 40-photo inspections at max camera size on
low-storage devices; chunk resume after hours offline; hash-mismatch and
corrupt-chunk recovery.
Security/compliance. Encrypted-store verification, token-grace
enforcement, purge-scope tests (bytes + thumbnails + caches gone; receipt
retained), purge-after-long-offline exception path, audit completeness
(every applied op has exactly one entry).
Field/pilot testing. Real basements with real dead zones; supervised
pilot with daily telemetry review; supervisor usability of conflict and
pending-change warnings.
Regression. Existing online-only flows unchanged when flags are off;
read-through-cache behavior preserved for unpinned inspections.
14. Milestone sequence with dependencies and acceptance criteria
No calendar dates. Each milestone lists dependencies and the acceptance bar
that gates the next. Order is intentional; M2–M4 can overlap in design but
land in sequence.
M1. Local durability (no sync yet).
Dependencies: decisions D1–D3 (draft model, conflict defaults,
photo policy) to avoid rework; none otherwise.
Acceptance: edits autosave to SQLite transactionally; kill/reboot loses
nothing; 40-photo queue + signature persist locally; purge-table schema
present but inert; unit + crash tests green.
M2. Pull + pin for offline.
Dependencies: M1; decision D5 (schema migration rule).
Acceptance: assignments pin with schema + version snapshot; "available
offline" and last-synced UI; stale-schema and unpinned-access handling;
integration tests for pin/refresh/evict green.
M3. Push with idempotent outbox + conflict v1.
Dependencies: M2; server idempotency + 409-with-state (decision D2
shapes the contract); decision D4 (accept-with-pending rule).
Acceptance: ordered push with resume; duplicate POSTs produce one effect
and one audit entry; conflict screen with per-field choice;
supervisor-authoritative denylist enforced; conflict matrix tests green;
no silent overwrites in either direction.
M4. Resumable attachments + submit barrier.
Dependencies: M3; decision D3 (original vs. compressed record).
Acceptance: chunk-resume upload with hash verification; 40-photo stress
passes on low-storage devices; submit blocked until required uploads ack
(per policy); signature binding + post-sign-edit rule enforced.
M5. Purge, audit, and security hardening.
Dependencies: M3–M4; decisions D6–D7 (offline auth grace, purge clock).
Acceptance: acceptance triggers purge scheduling; all customer bytes
(answers, photos, signature, thumbnails, caches) gone within the window
including after reboot/background; purge_log receipts retained;
long-offline purge-delay exception path tested; audit completeness
verified; encrypted-store and token-grace tests green.
M6. Observability + supervisor surfaces.
Dependencies: M3–M5 (events to observe).
Acceptance: device + server dashboards live; inspector sync states and
supervisor pending/conflict badges shipped; pilot alerts (purge breach,
sync stall, audit gap) firing on synthetic faults.
M7. Pilot.
Dependencies: M1–M6; pilot cohort and flag plan; decision D8 (pilot
success thresholds).
Acceptance: pilot entry checklist green (flags, rollback runbook,
support channel); sustained window meeting D8 thresholds; zero purge
breaches; supervisor sign-off; go/no-go for broad rollout recorded.
M8. Broad rollout + hardening.
Dependencies: M7 go decision.
Acceptance: progressive enablement with pause criteria monitored; rollback
drill executed at least once in production-like conditions; post-rollout
review captures conflict-rate, upload-completion, and purge-latency data
for follow-ups (e.g., widening offline scope).
15. Decisions required before implementation (with recommendations)
ID
Decision
Recommendation
Why
D1
Offline draft model: autosave-every-change vs. explicit save?
Autosave every change locally as the source of truth for drafts.
The brief's core complaint is discarded notes; explicit-save preserves the failure mode. Autosave + outbox is the standard fix.
D2
Conflict default: who wins true field conflicts, and which fields are supervisor-authoritative?
Inspector chooses per field; accept/reject + assignment always server-wins (denylist). Auto-merge non-overlapping and same-value writes.
Avoids silent loss without building a collaborative editor; keeps status transitions authoritative where safety demands it.
D3
Photo record: original bytes or defined compressed rendition?
Ask compliance; default to compressed rendition as the record (stated resolution/quality floor) if permitted, else originals.
40 originals over intermittent links is the main upload risk; a defined rendition sharply improves completion rates. Do not compress silently — record the policy.
D4
Supervisor accept/reject while tablet changes are pending: block, warn, or allow?
Warn + allow with "accepted with pending changes" state; block only final archival until tablet converges or supervisor overrides explicitly.
Blocking every accept on tablet sync stalls supervisors; silent allow risks accepting stale state. The middle path preserves velocity with visibility.
D5
Form-schema bump mid-offline: auto-migrate or force re-sync?
Auto-migrate additive changes; require guided re-sync for breaking changes (never silently drop answers).
Additive bumps are common and safe to migrate; breaking changes need human eyes.
D6
Offline auth grace: how long can a cached login authorize offline work?
Bounded grace (illustrative: one shift), then re-auth required; sensitive actions (signature, submit) require recent auth. Confirm exact bound with security/compliance.
Balances basement usability against lost-device risk; exact bound is a policy call, not an engineering guess.
D7
Purge clock: from server acceptance or from tablet observation?
Server acceptance starts the obligation; tablet purges within 24h of observation and reports delay as a compliance exception if observation was late. Confirm with legal/compliance.
Tablets cannot act on what they cannot see; the exception path makes the gap auditable instead of silent.
D8
Pilot success thresholds: what earns broad rollout?
Pre-commit numeric gates (e.g., sync success, attachment completion, conflict-resolution rate, purge compliance, crash-free sessions) agreed with product before M7.
Prevents moving goalposts; thresholds must exist before pilot data arrives. Values, not the practice, are the open item.
16. Risks (pointer)
The three riskiest assumptions — supervisor/inspector edit overlap, photo
upload under intermittent connectivity, and the 24-hour purge clock — are
analyzed in RESPONSE.md. This plan mitigates each (guided conflicts,
resumable uploads + compression decision, purge scheduler + exception path)
but cannot eliminate them without the D-decisions above.
muse/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.5 / 10graded blind as submission D
Very rigorous about correctness. It shows that a reconnect-triggered timer cannot meet the 24-hour rule and proposes an acceptance barrier: register every tablet copy, freeze, reconcile, purge, then accept. The reasoning is sound, but the design is heavy for a pilot. Acceptance can be blocked indefinitely, every stale write requires review even when the edits are disjoint, and any supervisor edit invalidates the customer signature. The dense prose makes it harder for a team to act on.
Strengths
Correct analysis of offline and powered-off devices and of cryptographic versus physical erasure
Detailed replay and idempotency semantics, including old acknowledgements arriving after newer feed versions
Extensive verification matrix that covers clock rollback, freeze races and lost purge acknowledgements
Weaknesses
Acceptance blocks on every registered tablet, so a lost device stalls the record
Record-level conflict review even for disjoint fields adds review load
Signature invalidation by supervisor edits needs the signer again, which is impractical
Prose is dense; some milestone outputs are bundled (M4)
Evidence the grader checked
Security section: 'Starting a 24-hour timer when acceptance is next observed violates the stated deadline'
D4 record-level optimistic concurrency with explicit three-way field review
Testing table row 'Acceptance requested while a tablet is offline or powered off for longer than 24 hours'
Objective checks
Files
gpt-6.1-sol/06-product-planning/RESPONSE.md
Three riskiest assumptions
Acceptance can wait for verified tablet cleanup (A2; D1/D2/D7). A tablet cannot learn about server acceptance without connectivity, and an app cannot execute a timed wipe while powered off. Purging on the next reconnect would not guarantee removal within 24 hours after acceptance. PLAN.md recommends a server acceptance barrier: register every tablet copy, reconcile its pending work, complete an approved purge, and acknowledge all copies before acceptance. The risk is operational delay from unreachable/lost tablets, plus an unresolved definition of compliant deletion on encrypted flash/backups. Product and security/compliance must approve the workflow and prove the deletion method before implementation. If immediate acceptance is mandatory, the alternative needs verified platform-enforced leases and explicit approval of cryptographic erasure; an app timer is insufficient.
Existing REST/version APIs can be extended into a reliable commit protocol (A3; D3). Integer record versions alone do not prevent duplicate effects after response loss or guarantee atomic audit history. The plan requires transactional version checks, mutation/idempotency/audit commits, all-writer compatibility, attachment verification, ordered change discovery, and a copy registry coordinated with acceptance. If these guarantees are absent, synchronization can overwrite supervisor work, repeat effects, omit audit events, or accept while untracked copies remain. Validate the contract in M0/M1 using duplicate, concurrent-write, server-restart, response-loss, and preparation-versus-acceptance tests. Do not compensate by dropping queued edits or blindly retrying expired operation IDs.
A signature binds a stable inspection snapshot and can be recollected after conflicts (A4; D4/D5). The brief does not define signed content or whether the signer will be available again after a supervisor edits the inspection. The recommendation binds the signature to canonical forms, notes, template, and the complete photo manifest; a changed signed field requires review and a new signature. If re-signing is infeasible or legal requirements differ, this may obstruct the field workflow and acceptance. Agree the signed projection, provenance, evidentiary format, and correction process before implementation, then test an offline signature followed by supervisor edits, attachment changes, conflict resolution, and re-signing. Never silently transfer a signature to changed evidence.
The implementation sequence, acceptance criteria, scope, local data model, sync and conflict contracts, attachments, security, migration/backfill, observability, staged rollout, rollback, and verification are in PLAN.md. Every requested planning deliverable is included; the blocking product/compliance/technical decisions are explicitly proposed rather than represented as already approved.
gpt-6.1-sol/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
gpt-6.1-sol/06-product-planning/PLAN.md
Offline inspection implementation plan
Basis and intended outcome
BRIEF.md is the only product input. This is a proposed implementation plan; recommendations and assumptions below require the named decisions, and are not claims about existing capabilities. No code, external sources, delivery dates, or staffing estimates are used.
The brief establishes these requirements:
Inspectors use company-managed tablets in places without signal; losing connectivity currently risks discarding notes.
Assigned inspections must support offline forms, up to 40 photos, and a signature.
Connectivity returns intermittently, and supervisors can edit the same inspection on the web.
Audit history is required, and local customer data must be removed within 24 hours after acceptance.
The server uses REST and integer record versions. Mobile SQLite is currently a read-through cache.
Product wants a pilot before broad rollout and excludes a full collaborative editor.
The outcome is durable local work with visible synchronization status, explicit conflict resolution, auditable server commits, and a verified cleanup path. A local save, a server synchronization, a submission, and acceptance are separate events.
Assumptions and scope
ID
Proposed assumption
Consequence and validation
A1
Inspectors can prepare assigned inspections while connected before entering a basement.
New assignments, first login, missing forms, and uncached inspections cannot be opened offline. Validate this workflow with the pilot; show readiness per inspection.
A2
Acceptance can wait until all registered tablet copies are reconciled and purged.
This makes the deletion deadline enforceable without assuming connectivity. Product must approve the acceptance delay; security/compliance must approve what constitutes deletion. This is a launch blocker.
A3
Server changes can provide atomic conditional writes, durable idempotency, audit records, change discovery, and a tablet-copy registry.
REST and integer versions alone do not provide these guarantees. Prove them in a server contract spike before depending on them.
A4
A signature attests to a defined inspection snapshot, and edits to that snapshot require a new signature.
Define the signed fields and evidence with product/compliance. Concurrent supervisor edits may require review and re-signing.
A5
A managed tablet profile belongs to one signed-in inspector at a time, with an approved offline authorization window.
Verify enrollment, lock behavior, credential expiry, account changes, supported OS versions, and key-storage capabilities.
A6
There is a bounded supported workload: 40 photos per inspection plus a signature, with capped photo size and a bounded number of prepared inspections.
Measure on pilot tablets. Decide caps before implementation; do not invent fleet size, traffic, or storage capacity.
A7
Accepted inspections become immutable for tablet editing.
Corrections create a separately audited revision and a new signing/acceptance cycle. Confirm existing web workflows can honor this.
In scope: preparation/download, durable forms and notes, camera capture, signature capture, offline reopen, background and foreground synchronization, conflict review, submission, acceptance cleanup, migration, support diagnostics, and a controlled pilot. Web changes needed for version checks, conflict visibility, signature invalidation, and acceptance are included.
Out of scope: live cursors or collaborative text editing, offline creation of unassigned inspections, offline supervisor acceptance, arbitrary customer-data exports, and automatic legal interpretations. Device recovery through an authorized support workflow is included; a general offline sharing feature is not.
Reliability target: once the UI says “Saved on tablet,” a crash, restart, or connectivity change must preserve the saved edit. Storage failure must stop the saved indicator. Performance targets for local save latency, opening a prepared inspection, text sync, and draining a 40-photo inspection will be agreed from representative-device measurements in M0. Ownership is by role: mobile owns durable storage/workers; backend owns commit/audit/acceptance contracts; product owns workflow; security/compliance owns retention and signatures; operations owns alerts and recovery procedures. Roles are responsibilities, not staffing estimates.
Decisions required before implementation
All choices are provisional recommendations. M0 records the accepted choice, decision owner, evidence, and unresolved constraints.
ID / owner
Decision and recommendation
Alternative and tradeoff
D1 / product + compliance
Define “accepted” as a server transition after validation and confirmed cleanup of every registered tablet copy. Keep submission distinct; label the intermediate state “Awaiting tablet cleanup.”
Acceptance before cleanup requires a proven deletion mechanism independent of reconnecting; a reconnect-triggered 24-hour timer cannot meet the brief.
D2 / security + compliance + mobile
Approve the deletion standard, all locations containing customer data, and platform capabilities. Prefer encrypted per-inspection payloads/files and verified deletion of their keys, files, database rows, and transient copies, with managed backups excluded.
Literal physical erasure of flash remnants or timed deletion on an offline, powered-off device cannot be promised by an app. If cryptographic erasure is unacceptable, obtain an implementable approved method before proceeding.
D3 / backend
Add a REST sync contract with compare-and-swap versions, transactional audit and idempotency, attachment staging, paged changes, copy registration, and acceptance coordination.
Existing CRUD without these guarantees risks duplicate writes and missing audit events. A complete event-sourced backend would add unnecessary scope.
D4 / product + backend
Use record-level optimistic concurrency and explicit three-way field review for conflicts; never silently overwrite either author's work.
Automatic disjoint-field merging reduces reviews but adds semantic ambiguity. Defer it until pilot evidence justifies it; do not build a collaborative editor.
D5 / product + compliance
Define the signed projection, signature provenance, and evidence retention. Recommend signing forms, notes, template version, and the complete photo manifest using an immutable snapshot hash. A changed signed field invalidates the signature.
Excluding some content is possible only with explicit approval; silently carrying signatures across changed evidence is unacceptable.
D6 / product + mobile + security
Set maximum photo bytes/resolution, prepared-inspection count, offline authorization duration, and disk reserve from field needs and tablet measurements. Count 40 photos across the inspection; the signature is a separate artifact.
An unlimited offline window or unlimited prefetch increases authorization and storage risks. Shorter windows must still cover the agreed basement workflow.
D7 / product + backend
Freeze editing during final acceptance coordination, reconcile every tablet, then purge and accept. Cancellation/revision requires explicit server action and preparation again.
Offline tablets can delay acceptance. If that is unacceptable, redesign D1/D2 before building; merely removing devices from the registry is not proof of deletion.
D8 / product + security
On expired authorization, lock local data and pause upload until reauthentication; on reassignment/revocation, require supervised recovery or approved disposal.
Automatic upload with stale permissions exposes data; automatic destruction may lose unsynced work. Define the authorized recovery and disposal policy before pilot.
D9 / compliance + backend
Store authoritative audit history on the server with policy-controlled access and retention. Preserve actor, before/after values, versions, signature evidence, and resolution provenance.
The brief provides no audit retention duration or signature format; do not invent either. Define whether offline per-edit provenance is required and its evidentiary limits.
D10 / product + operations
Specify pilot workload, success thresholds, stop conditions, and support ownership before exposure. Require zero known saved-edit losses, duplicate committed effects, unsigned acceptance, or cleanup-policy violations.
Expansion based only on enrollment count or connectivity averages misses correctness failures. Quantitative performance/support thresholds need M0 measurements.
D11 / product + mobile + backend
Pin the prepared template for the draft and continue accepting it while that template remains approved. If it is retired, require an explicit field-mapping review and re-signing before submission.
Silently swapping forms can discard answers or change signed meaning. Define compatibility and retirement rules before implementation.
Inspector and supervisor flow
Prepare while connected. Authenticate and validate assignment. Register the device copy before downloading any customer data; pin the form/template version, reference data, and current server version. Check capacity and download atomically. “Ready offline” appears only after the complete inspection and form dependencies are locally durable. Incomplete downloads remain unavailable offline. Show last synchronization and offline-access expiry.
Work offline. Open a prepared inspection, fill forms/notes, and capture photos. Save each meaningful edit locally without waiting for HTTP. Photos become saved only after encrypted bytes and a durable manifest entry exist. Show “Saved on tablet” separately from “Synced.” At the photo/storage limit, explain the limit and preserve existing work. Reopen the same draft after an app or tablet restart.
Sign and submit. Validate the local form, freeze a signed snapshot, and capture its signature offline. Submission offline is an intent marked “Waiting to submit.” It is not server submission or acceptance. Editing signed content explicitly invalidates the signature and requires re-signing. The UI must never show an incomplete signature/blob dependency as submitted.
Reconnect intermittently. The worker resumes automatically on foreground/network opportunities. Notes need not wait for large photo transfers. Display pending items, progress, retryable failure, sign-in required, and conflict separately. Interruptions do not discard the draft, reset completed chunks, or duplicate mutations.
Resolve concurrent edits. A supervisor's web edits increment versions and enter the audit log. A version mismatch leaves both versions available for review. Compare the last common base, the tablet changes, and the current server values; choose values or rewrite deliberately. Conflict resolution creates a new mutation against the current version. Review any changed signed content and collect a new signature as necessary.
Accept and clean up. A supervisor requests acceptance of a submitted candidate. The server freezes new editing/preparation and requests reconciliation and cleanup from every registered tablet copy. An offline tablet keeps the request pending. Reconnect brings its saved work into review, including any work made before it learned of the freeze. Resolve that work before a tablet discards it. Once all copies have been purged and acknowledged and the complete candidate is valid, the server atomically accepts. The tablet shows a data-free receipt only if its fields are approved as noncustomer data; accepted content cannot be downloaded again to tablets.
Local data model and durability
Use SQLite as the durable draft store, not an evictable cache. Encrypt customer-bearing payloads using an approved per-inspection key scheme; encrypt attachment files too. Verify database/WAL/journal/temp and backup handling in D2. Never include inspection data in logs or crash reports.
Local entity
Important fields and behavior
Inspection baseline
Stable server inspection ID, tenant/user scope, confirmed integer version, assignment and lifecycle state, pinned template, confirmed server payload. Payload is encrypted.
Draft/revision
Client-generated revision ID, baseline reference, local sequence, dirty fields and encrypted values, local validation, immutable signed/submitted snapshot hash when applicable. Maintain pending local edits separately from confirmed server data.
Outbox operation
Globally unique operation ID, actor and installation ID, local sequence, baseline version/snapshot, immutable patch, dependencies, persisted request fingerprint once sent, retry state, and acknowledgement. Encrypt customer-bearing request bodies.
Attachment
Client-generated ID, kind (photo/signature), encrypted file reference, byte length, content hash, upload/session/chunk state, owning revision, and server verification state. A signature also references its signed snapshot and declared signer.
Sync checkpoint
Scoped feed cursor, last confirmed server contact, schema/contract version, and worker state. It must advance in the same transaction as applying a change page.
Copy/cleanup state
Registered copy token/generation, freeze/cleanup request state, purge progress. Retain only approved noncustomer recovery metadata after purge. The authoritative copy registry is on the server.
Persist a form edit, its draft revision, and its outbox/provenance entry in one transaction before acknowledging local save. Use durable transaction settings and measure their latency. For files, write encrypted temporary bytes, flush/close, atomically move to the private destination, then commit the manifest/outbox. Recovery sweeps orphan temporary/files and identifies missing manifests without deleting referenced unsynced evidence. A captured photo cannot be declared saved before both pieces are durable.
Serialize workers per inspection. Client UUIDs identify revisions, operations, and attachments without offline ID collisions; preserve existing server inspection IDs without assuming their format. Persist signed snapshots separately from the mutable draft. Compaction may batch unsent operations only if the required per-edit provenance and signing boundaries survive; never rewrite a request already sent. Database/schema upgrades are transactional and preserve queued work. Insufficient disk space must produce an actionable save failure, with the previous durable version intact.
Server and synchronization contract
Extend existing REST APIs additively. The following describes required semantics, not existing routes.
Preparation and changes. An authorized preparation request transactionally registers a device copy and returns a consistent snapshot/version, template requirements, and scoped synchronization cursor. Do not release content before registration. A paged changes API uses an opaque ordered cursor, not per-record integer versions as a global sequence. Include edits, assignment/permission changes, acceptance/freeze state, and deletion tombstones. Every tablet read path, including feeds, refreshes, and legacy endpoints, must enforce registration/freeze rules and avoid returning accepted customer payloads; acceptance notifications carry only approved control metadata. Define cursor expiry: an expired cursor triggers a scoped snapshot refresh while preserving the outbox and its original bases.
Mutation. Send operation ID, immutable request fingerprint, expected integer version, actor/device context, schema/template version, and a patch. The server checks authorization and lifecycle, then atomically commits the patch, incremented version, audit event, and idempotency result. Every writer, including the web app, uses the same version discipline. Return the committed operation/result/version. A version mismatch returns the current version and conflict information without changing inspection data.
Replay. Repeating an authenticated operation ID with the identical request returns the original result, even if the inspection has since advanced. Reusing it for a different body is an error. Enforce uniqueness transactionally across worker restarts and concurrent requests. Keep deduplication records for the supported retry/recovery lifetime; if records must expire, reject old operations and provide a result/status reconciliation path instead of blindly applying them again.
Client scheduling. Upload independent binaries and pull changes as connectivity permits, then serialize inspection commits respecting dependencies. For multiple offline edits, compute each expected server version from the acknowledged predecessor; persist the exact request before first transmission. Do not guess successive versions when other writers exist. On timeout, retry the same request. Acknowledgement updates the confirmed baseline and outbox atomically while retaining newer unsent overlays. If a replay response is older than an already-pulled version, mark the operation acknowledged without regressing the baseline; refresh current state and review remaining operations against their retained bases. A crash after server commit but before local acknowledgement is safe.
Error handling. Back off transient/network/server errors with jitter and a bounded retry interval; honor server throttling. Unauthorized requests pause for sign-in, forbidden access locks work for recovery, conflicts require review, and validation/schema errors require correction. Cancellation of an HTTP attempt does not mean the server rolled it back. Do not classify permanent errors as successful or discard their payloads.
Audit. The server commit time is authoritative. Retain claimed offline edit time separately, with actor, device, operation ID, before/after or an equivalent reconstructable revision, versions, conflicts/resolutions, photo/signature hashes, submission, acceptance, and cleanup acknowledgements. Offline timestamps/claimed authorship are not trusted proof of actual time. Audit writes failing means the corresponding mutation fails. Restrict audit alteration/access and define retention under D9.
Recommended lifecycle: working draft → signed snapshot → submitted candidate → acceptance pending cleanup → accepted. Validation failure, conflict, or changed signed content returns the candidate to review and re-signing as necessary; it never becomes accepted automatically.
Conflict and signature policy
Reject a stale record write even if the apparent changes are disjoint. Keep the encrypted local base and draft while retrieving the current authorized server version; show field-level differences and the author of server changes. Present long notes as whole-field alternatives rather than attempting a text editor merge. Resolving disjoint changes is a deliberate review in v1. If the common base is unavailable, require full comparison instead of inferring an ancestor.
Resolve or supersede the blocked operation using a new operation ID and current expected version; never change the body under an existing idempotency key. Archive the reviewed base/local/server alternatives and the chosen disposition as authorized server audit evidence before considering reconciliation complete. Every saved operation must be accounted for as committed or explicitly resolved; no unresolved branch may be purged just because another branch was submitted. A new server edit during review conflicts again.
Deletion, reassignment, freeze, template changes, and accepted state are lifecycle conflicts, not ordinary field conflicts. Stop the normal queue and invoke the relevant recovery/cleanup workflow. A freeze may be unknown to an offline tablet; retain its later saved edits for review when it reconnects. Reconciliation during acceptance uses explicitly authorized/versioned requests, never the ordinary edit path to bypass the freeze.
Photos have stable IDs; a retry is not another photo. Concurrent additions/removals still change the inspection manifest and require version review. Enforce the total 40-photo limit server-side across all writers and client-side for the local draft; excess concurrent additions remain pending for explicit selection, not silent dropping. Removing evidence retains its audit/hash and any server retention required by policy.
The signature records the signed projection's hash, template version, declared signer, capturing inspector/device, and capture-time provenance. Canonicalization must be specified and identical on tablet/server. Changes to any included form field, note, or attachment manifest invalidate the signature, including supervisor edits. The server recomputes and verifies the projection before submission/acceptance. Keeping an old signature image does not make it valid for a new revision. Signature capture is evidence collection; its legal adequacy is a D5 decision.
Attachments
Capture into application-private encrypted storage; prevent gallery copies, thumbnail leakage, and backup copies. Include the signature in this protection. Decide evidentiary image format/metadata before capture; any approved resizing or metadata removal occurs before the immutable content hash and signature, and must not silently alter signed evidence.
Use resumable, authenticated uploads with a client attachment ID, declared length/hash, and acknowledged chunk offsets. Refresh expired upload credentials after reauthentication; never log upload URLs. Verify complete bytes/hash server-side before returning a verified blob reference. Uploading/finalizing binary staging does not by itself attach evidence to an inspection; the versioned manifest mutation does. Identical retries return the same object state. Reconcile expired sessions rather than counting a replacement upload as another photo.
Dependencies require verified photo/signature objects before their manifest/signature operations, and all signed content before submission. Text commits can proceed independently before a slow photo upload; the candidate remains incomplete until all dependencies match its frozen snapshot. Edits during upload produce a new revision and invalidate any affected signing intent. Staged unreferenced uploads receive a bounded server cleanup policy; uploaded evidence referenced by an accepted candidate is never treated as an orphan. Local files are retained until verified/reconciled and the approved purge workflow authorizes removal.
Capacity admission includes 40 photos at the chosen maximum size, signature, encryption overhead, temporary capture space, migration space, and all other prepared inspections. Enforce the configured reserve during capture, not only at preparation. Measure battery, resume behavior, and upload costs on real pilot hardware; tune chunk sizes/concurrency from those results.
Security, acceptance, and the 24-hour requirement
Use authenticated transport, current server authorization on every operation, managed-device enrollment, approved at-rest encryption and key storage, and an expiring offline authorization grant. Bind stored data and outbox work to tenant, account, and installation. Lock on credential expiry/background timeout according to policy. Account switching must not expose another inspector's draft; pending work needs authorized drain/recovery before destructive logout. Device loss triggers managed response and registry reconciliation, but issuing a wipe command is not proof it executed.
An ordinary app cannot react to acceptance it has not learned of or execute a wipe while powered off. Starting a 24-hour timer when acceptance is next observed violates the stated deadline if acceptance happened earlier. Cryptographic erasure and physical deletion are also different standards. These are D1/D2 launch decisions, not details to defer to QA.
The recommended v1 protocol is an acceptance barrier:
Every delivery of tablet customer data registers the copy before content/key release; registration and acceptance-freeze acquisition serialize to prevent a download race. No accepted inspection is available for tablet preparation, and a frozen inspection cannot acquire a new copy.
A supervisor requests acceptance at a specific version. The server validates it and freezes ordinary edits. Track all existing copies, including managed copies left by migrations/reinstalls; never remove one merely because it is unreachable.
Each tablet reconnects, stops new edits, clears decrypted in-memory views, and reconciles all saved operations and attachment dependencies. Late offline work is reviewed explicitly. If the candidate changes, its signature is invalidated and the acceptance request is revised or canceled for re-signing; already purged tablets prepare again if further work is needed.
After review is complete, the tablet performs an idempotent, recoverable purge: inspection rows and payloads, outbox bodies, attachment/signature/thumbnail/temp files, keys, in-memory objects, and classified customer-bearing sync metadata. The approved method must address SQLite WAL/journal/free pages and OS caches/backups, not just DELETE statements. Persist only approved noncustomer cleanup progress. A crash resumes purge before displaying customer data.
Acknowledge completed purge to the server for the specific copy generation and cleanup job. Repreparation creates a new generation, so an old acknowledgement cannot clear a new copy. Lost acknowledgements are recoverable by repeating the server's cleanup job and confirming already-absent material; acknowledgements need not retain customer data on the tablet. The server atomically verifies every copy's acknowledgement, complete evidence, current version, and valid signature, then records acceptance and its audit event. Thus approved deletion precedes acceptance, satisfying the 24-hour upper bound with margin.
Keep the acceptance freeze until completion or explicit cancellation; deletion errors leave acceptance pending and alert operations. Unreachable tablets visibly block acceptance. New data access after purge requires fresh authorized preparation, and accepted records remain blocked. A lost tablet requires verifiable approved erasure or acceptance remains blocked. The server retains the authoritative inspection/evidence/audit according to the approved server retention policy.
If product requires immediate acceptance despite disconnected copies, the recommendation must change before implementation: evaluate a platform-enforced per-inspection access/key lease expiring within 24 hours of a server assertion that the record is still unaccepted, with clock-rollback resistance and compliance-approved cryptographic erasure. A lease would bound offline usability and still cannot promise physical deletion while powered off. An app timer, wall-clock comparison at next launch, or an unverified hardware capability is insufficient. No pilot acceptance path may bypass D1/D2.
Migration and backfill
Deploy additive server contracts first, with compatibility for the existing client. Inventory only the legacy mobile cache/data behaviors needed for this feature during implementation; this plan assumes no repository investigation. Add the draft/outbox/attachment/checkpoint schema using a transactional, restartable migration. Version the schema and contract separately; reject an unsupported downgrade before it touches queued work.
Treat clean legacy cached records as baselines needing online authorization, fresh version/template validation, and copy registration before “Ready offline.” They are not proof of user intent. Preserve detectable unsynced legacy edits in encrypted recovery storage and require reconciliation; if their base version is unknown, do not fabricate it or upload as a fresh overwrite. Explicitly report any legacy work that cannot be recovered; this change cannot reconstruct notes already discarded before upgrade.
Backfill stable attachment identities/manifests and an audit baseline from actual existing server records/version metadata. Mark the baseline as a migration event; do not invent past edit history. Capture new authoritative audit events from the cutover onward for web and mobile. If historical retention obligations need more, D9 must supply a real source or resolve the gap before pilot.
Register existing tablet customer copies during enrollment, or purge and reprepare them before acceptance-enabled pilot use. Include old cache databases/files and existing backup policy in cleanup verification. Keep schemas additive through rollout, and verify old/new clients cannot bypass freeze, version, or acceptance rules. A successful upgrade preserves pending work and hashes; an interrupted upgrade restarts safely with enough reserved disk space.
Observability and operational recovery
Track local durable-save failures, oldest pending operation age, pending bytes/count, last successful synchronization, retries/error classes, idempotency replay counts, conflict incidence/age, attachment checksum/resume failures, signature invalidations, submission latency, freeze duration, registered outstanding copies, purge failures, and acceptance-to-cleanup invariant violations. Distinguish expected offline queue age from failed synchronization after connectivity returns.
Use scoped opaque operation/correlation IDs and aggregated metrics. Do not log notes, customer names, signature images, photo contents, credentials, or raw payloads; review even IDs for customer association. Export diagnostics only through the authorized support path. Delayed telemetry from offline tablets means absence of an alert is not evidence of successful deletion; the acceptance registry is the enforcement mechanism.
Alert on accepted records without all cleanup acknowledgements, missing audit transactions, checksum mismatches, unrecoverable saved data, and stuck acceptance/purge jobs. Agree queue-age, retry, conflict, and performance thresholds in M0. Provide runbooks for reauthentication, version conflicts, failed uploads, low disk, expired cursors, interrupted migrations/purge, lost tablets, and pending acceptance. Operators can inspect server status and authorized provenance; they cannot silently overwrite or mark a wipe complete.
Testing and verification
Run contract, migration, security, and device fault tests before field exposure. Validate invariants at each crash boundary, not just end-to-end success.
Scenario
Required result
Lose signal during editing; kill/restart after local transaction acknowledgement
Every acknowledged edit returns; UI never reported unsaved work as saved.
Crash before/after photo file move or manifest commit; disk fills
Saved photos remain hash-identical; incomplete capture is reported; orphan recovery does not delete referenced evidence.
Server commits, response disappears; client/server restart; duplicate workers
Retry applies one effect and one commit audit event; original response can be reconciled; newer local edits survive acknowledgement.
An old replay acknowledgement arrives after a newer feed version
The operation is acknowledged once, the baseline never regresses, and remaining drafts keep their bases for review.
Supervisor and tablet edit same or different fields from one version
Stale write fails; base/local/server comparison is preserved; reviewed resolution is versioned/audited; a second race conflicts again.
Notes/photos/template change after signing; client/server canonicalization differs
Invalid signature cannot submit/accept; valid identical snapshots match; re-signing binds the new evidence.
No skipped committed change; checkpoint/application are atomic; refresh preserves local drafts; lifecycle/permission changes are enforced.
Token expires, account/tenant changes, tablet is revoked
Data is locked/scoped; no stale-authority upload or cross-account exposure; approved recovery is usable.
Acceptance requested while a tablet is offline or powered off for longer than 24 hours
Server stays pending; it cannot accept without confirmed approved purge. Reconnect preserves late work for review before cleanup.
One tablet purges, another reports late edits; acceptance is revised/canceled
Reviewed alternatives are archived, changed evidence requires re-signing, purged tablets must reprepare for new work, and no stale candidate is accepted.
Preparation races acceptance freeze; stale web client submits edits
Copy registration/freeze ordering is consistent; neither new copy delivery nor old-client writes bypass the barrier.
Crash during purge, acknowledgement lost, tablet storage/key deletion fails
Purge resumes before exposure, acknowledgements are idempotent, failure blocks acceptance, all storage locations meet D2.
Device clock rolls backward/forward; telemetry never arrives
Server time drives acceptance/audit; no client timer creates a false compliance claim. Test clock resistance if the lease alternative is selected.
Upgrade mixed/partial legacy caches with unsynced work; interrupt migration
Work is preserved or explicitly quarantined; no invented base/history; stale clean data is not called offline-ready.
Disable pilot flags with queued notes/photos/conflicts and pending cleanup
Workers, conflict recovery, server contracts, and purge continue; no downgrade or cache eviction destroys work.
Representative tablets under full permitted load and poor/flapping connectivity
Local latency, memory/disk, battery, resume behavior, and recovery meet the M0 workload and performance criteria.
Use two-client/web concurrency fixtures and fault injection at every local/server transaction and attachment phase. Security verification inspects files, database sidecars, OS capture paths, backups, memory/session clearing, and redacted diagnostics against the approved deletion standard. Pilot UAT walks preparation → basement work → restart → intermittent sync → supervisor conflict → re-sign → submission → cleanup → acceptance, including a blocked-acceptance case.
Milestones, dependencies, and acceptance criteria
No dates or headcounts are assumed. Completion is evidence-based; implementation starts only after M0's blocking decisions.
Milestone
Dependencies
Reviewable output and acceptance criteria
M0 — Decide and prove feasibility
None
Decision log for D1–D11; sample user flows, sync/audit contract, supported workload/performance thresholds, pilot/stop criteria. Discovery prototypes prove durable storage, signature canonicalization, and approved deletion on supported tablets. Product accepts offline preparation and acceptance delay. Unproven D1/D2/A3 blocks feature implementation.
M1 — Server foundations
M0
Additive version/idempotency/audit, feed/tombstones, blob staging, registry/freeze/acceptance contracts. Duplicate/timeout/concurrent-writer tests pass; audit failure aborts commit; old web/mobile paths cannot bypass safety checks. No acceptance without all purge acknowledgements and valid evidence.
M2 — Durable tablet work
M0; stable M1 contract
Encrypted baseline/draft/outbox/file store, restartable migration, authorization lock, preparation readiness and local-save UI. Crash/disk/upgrade tests preserve acknowledged notes and evidence; uncached records stay unavailable offline. M1/M2 work may overlap only after contracts are agreed.
M3 — Synchronization and conflicts
M1 + M2
Resumable worker, paged pull, retry/idempotent acknowledgement, three-way conflict UI and web compatibility. Intermittent connectivity and response-loss tests pass with no duplicate effects or overwritten drafts. Review/re-sign behavior is understandable in UAT.
M4 — Evidence, submission, cleanup
M3; verified blob foundation from M1
40-photo/signature flow, snapshot verification, upload dependencies, frozen acceptance reconciliation, full purge and acknowledgements. A complete offline inspection can become accepted; disconnected or failed-cleanup copies keep it pending. Security/compliance verifies D2 against actual storage.
M5 — Operational readiness
M4
Fault suite, mixed-version migration/rollback rehearsal, privacy-reviewed metrics/alerts, runbooks, and support recovery pass. No unresolved safety invariant failure; agreed performance tests pass on pilot tablets. Rehearse a disablement while queues and acceptance jobs are active.
M6 — Controlled pilot
M5
Enroll an explicitly allowlisted cohort and bounded assignments/devices. All are prepared and registered; exercise real basement and supervisor-conflict flows. Review each acceptance and audit/cleanup evidence. Meet D10 thresholds over the agreed observation/workload window, with no known safety failures.
M7 — Expand and maintain
M6 passes
Expand cohort in stages with health review at each stage; pause on regressions. Keep migration/compatibility and retention operations supported until old queues/copies are resolved. Broad rollout requires product, engineering, and compliance signoff on measured pilot results.
Staged rollout and rollback
Roll out server additions and protections first, then instrumented internal tablets, then the allowlisted pilot, then increasing cohorts. Gate preparation/offline capture separately from sync/cleanup support. For every enabled inspector, validate assignment, enrollment, free space, registered copies, and offline readiness before entering the field. Do not classify acceptance-pending cleanup as accepted in dashboards or user-facing messages.
Pause expansion and stop new offline enrollment/preparation if saved data is lost, audit or idempotency invariants fail, signatures bind incorrectly, authorization leaks occur, deletion cannot be verified, or agreed performance/support thresholds are exceeded. Investigate queued/conflicted work with the runbooks; retain it encrypted until resolved under approved policy.
Rollback stops new exposure while keeping installed-client upload, conflict resolution, blob resumption, audit, registry, freeze, and purge paths alive. Do not disable cleanup or acceptance protections. Keep additive server schema/contract support; drain or explicitly recover every queue before any database downgrade/removal. Do not revert tablets to a cache-only app while drafts exist. For a severe security incident, lock affected data and use the approved managed response, recording any authorized disposal rather than silently discarding work.
Resume rollout only after the root cause is fixed, the relevant fault tests pass, affected drafts/copies are accounted for, and M6's review gates pass again. Completion means a supported offline-to-acceptance flow with audit and cleanup evidence, not merely a successful upload on a connected tablet.
gpt-6.1-sol/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.5 / 10graded blind as submission H
The most implementable protocol of the set: per-field modified_version merge on the existing integer version, exact per-op result semantics, seq-gap handling, and well-reasoned resolution rewrites. The purge design, however, depends on redefining 'accepted' as the server committing the inspector's submit, which removes supervisor acceptance. That is a large product reinterpretation, flagged only for compliance. Many specific numbers are set as defaults without justification (12-hour window, 3 MB, 50 ops, 45 days, 2%).
PLAN.md is buildable on the recommendations as written. These three assumptions are the ones that, if wrong, change the design rather than a parameter. They are A1–A3 there. Compliance still has to confirm A1, and the M1 API inventory has to confirm A2, before real customer data is piloted.
Acceptance and the 24-hour clock. The plan treats "accepted" as the server transaction that commits this tablet's signed submit, and treats the deletion rule as satisfied when the tablet erases customer content as soon as it learns that timestamp — usually in the same response, and always against server time. Local data is allowed to sit on the device before that event, including through a basement visit and a queued photo upload. That reading matches the brief's wording (removal after acceptance, not after download). It is unsafe if acceptance is actually a supervisor action that can happen while the tablet is offline, or if the duty is a wall-clock 24 hours from a server timestamp whether or not the device is powered on and reachable. In those readings a correct tablet can still be late, because a basement has no signal and a dead tablet cannot delete anything. The fallback in the plan — encryption, MDM remote wipe, and a compliance disposition for late purges — is weaker than a strict reading. If counsel requires the strict reading, the pilot cannot promise it; the product would need a different control than offline storage.
Field-level versions on today's integer REST records. Supervisors edit the same inspection on the web, and the product does not want a collaborative editor. The plan therefore merges only when the two sides touched different fields, and asks the inspector to pick a winner when they touched the same one. That works only if every web save bumps the existing integer version and records which fields changed. A single version on the whole document, or a write path that stores a blob without a diff, cannot tell an overlapping note from a harmless edit. The offline queue would then either block on every supervisor save or overwrite one side. Both outcomes throw away the reason for the project or violate the collaborative-editor constraint. M1 has to inventory the real write paths before the sync client is finished. If the web app cannot be brought through one versioned domain service, stop and re-scope; do not paper over it with last-write-wins.
Upload-before-accept on intermittent links. The server accepts, and the tablet then deletes, only after the form, the signature, and all photos (up to 40) are durable server-side. That avoids deleting the only copy before upload, and it avoids marking accepted a package the server cannot read. It also means a tablet that has left the basement and still lacks a usable link keeps the full customer data set, and the inspection is not yet accepted, so the 24-hour clock has not started. The plan assumes pilot devices normally finish that upload within the same shift, and that this delay is what the regulation allows. If signatures are supposed to count as acceptance immediately, or if sites routinely stay offline for days, the same rule cannot both protect unsent photos and clear the tablet on time. The outbox-age alert shows the condition; it does not remove it.
grok-4.7/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
grok-4.7/06-product-planning/PLAN.md
Offline inspection sync
This plan makes assigned building inspections editable on company-managed tablets with no signal. The inspector can open a prefetched inspection, fill the form, take up to 40 photos, and collect a signature. Every edit commits to SQLite before any network call. When connectivity returns, even briefly, the tablet uploads an idempotent operation log and the photo and signature bytes to the existing versioned REST records. Non-overlapping edits from a supervisor on the web merge automatically. Overlapping fields pause that inspection and the inspector chooses the value. After the server durably accepts the signed submission, the tablet deletes local customer data.
The current failure (a dropped connection after open discards notes) is closed by the same path: the form becomes editable only after a local snapshot commit, and online use also writes locally first. Sync is not a second editor. There is no real-time co-editing.
This plan has no calendar dates and no staffing estimates. Order, dependencies, and exit criteria are the commitments. Adopting the plan adopts the recommendations below. Two items still need an explicit compliance confirmation before real customer data is piloted: the meaning of "accepted," and what the server audit record stores. Those are D1 and D9.
Assumptions
Beliefs this plan depends on. A1–A3 are the ones most likely to force a redesign; they are summarized in RESPONSE.md.
ID
Assumption
A1
"Accepted" is the server transaction that commits this inspector's submit. Local customer data may remain on the tablet before that event. The 24-hour deletion duty is met by deleting as soon as the tablet learns of acceptance (normally in that same response), measured against server time. A tablet that never comes back online after a lost accept-response is covered by encryption plus remote wipe, which may be weaker than a strict wall-clock reading.
A2
Inspection writes can be given a per-field modified version under the integer version the APIs already use. The web app saves a form the server can diff. Supervisors and inspectors overlap on individual fields often enough to matter and rarely enough that a manual resolve step is acceptable. The product constraint against a collaborative editor means we will not build character-level merge, presence, or live co-editing.
A3
After leaving the basement, a pilot tablet normally gets enough connectivity the same shift to upload up to 40 photos and complete submit. Acceptance waits for those bytes, so customer data stays on the tablet until upload finishes. That delay is acceptable under A1.
A4
Today's SQLite data is a read-through cache. It holds no unique notes. Upgrade may discard it. Notes already lost to disconnects before this ships are unrecoverable.
A5
Forms are structured. Fields have stable ids. The pilot type set is text, long text, number, single select, multi select, boolean, date, plus inspection-level photos and one signature.
A6
"A signature" is one inspector signature per inspection: a raster image, the signer display name, and a client timestamp. It is required to submit. It is not one of the 40 photos.
A7
Assignments are created on the server. The tablet does not create inspections. The pilot is one inspector account per tablet. Unsupported extra writers (a second tablet) fall through the same conflict rules.
A8
Company mobile-device management can require disk encryption, a screen lock, remote wipe, and no backup of the app sandbox. The OS photo library is not a storage location for captures.
A9
The server keeps the regulatory audit history, including field values, on the server retention schedule. The 24-hour rule in the brief applies to data on the tablet.
A10
Refusing to open an inspection that has no local snapshot is acceptable. A start-of-shift prefetch while the tablet still has signal is part of the job.
A11
Auth is token-based. A bounded offline session can be added without a new identity system.
A12
Clocks on tablets drift. Server time is authoritative for acceptance, purge, and audit order. Client timestamps are claims.
A13
A feature flag can select pilot inspectors and devices. Shipping this does not require a new distribution channel.
A14
The threat model is a lost or stolen tablet, a stale or replayed write, and ordinary network failure. The basement network is not treated as a separate hostile environment beyond TLS.
Discovery at the start of milestone M1, from the current API rather than from this brief: which inspection resources exist, which of them already carry integer versions, which web write paths bump those versions, how photos are stored today, and how long tokens live. Where the code disagrees with A2 or A4, stop and revise this plan before pilot.
Decisions required before implementation
Each recommendation is the option to build. Compliance must confirm D1 and D9 before real customer data is enabled. The other rows are explicit so they are not relitigated mid-build; an objection belongs in the M0 record.
ID
Decision
Recommendation
Confirmation
D1
What "accepted" means, and when the 24-hour local-deletion clock starts
The only transition into accepted is the server committing the assignee's submit in the same transaction as validation and the audit row. accepted_at is that server timestamp. Supervisor web saves may change fields while status is assigned or in_progress; they do not accept the inspection. The tablet purges customer content immediately when it has durably stored a submit result (including a duplicate retry) that says accepted. Startup finishes a purge that was already authorized by a stored accept result and interrupted. Startup does not purge merely because a submit request might have reached the server; that case syncs first and purges only from the stored or retried result. Withdraw submit is allowed only while the submit op is still pending and has never been sent. Once it is in flight, the tablet waits for the result. After accept, fixes happen on the server and are audited there; the tablet does not download the content again. Immediate purge is the default. A short on-device review window is allowed only if it ends by accepted_at + 24 hours in server time.
Compliance and product, before real customer data
D2
How supervisor and inspector edits combine
Field-level optimistic concurrency on one integer inspection version. Non-overlapping fields both commit. Overlapping fields pause the inspection and the inspector resolves on the tablet. Photos merge by client-generated id. No collaborative editor, no last-write-wins of the whole inspection, no silent drop of either side.
Product, before M3
D3
How offline writes talk to REST
Add one idempotent sync request for the operation log and a separate resumable blob upload. Keep integer versions and the existing web REST writes. Both paths call one domain service that bumps the version and writes the audit row.
Adopting this plan
D4
How long an inspector can work after the last online authentication
Editing of already prefetched inspections remains available for 12 hours after the last successful online authentication, then the app locks. Unlock to keep editing uses the device credential (PIN or biometric) and does not require the network. Sync requires a live token. Lock, sign-out, and token expiry do not delete unsent work.
Security and product; 12 hours is the default
D5
How local data is protected and erased
Hardware-backed keystore holds a device key that wraps a per-inspection data key. Field values, header, photos, and signature are encrypted with that inspection key. Purge destroys the key and deletes the files and rows. MDM full-disk encryption remains mandatory and is not the only control.
Security, before pilot
D6
Schema changes while a tablet is offline
Schema versions are immutable. An inspection stores the schema version it was prefetched with and keeps that version through accept. Newer schemas apply to inspections prefetched afterward.
Adopting this plan
D7
Photo size
The cap of 40 is a hard client and server limit on non-deleted photos, signature excluded. Default capture processing: long edge 2048 px, JPEG quality 0.7, refuse a photo that is still over 3 MB after that. Refuse new captures when free space is under a configured floor; set the floor from the pilot tablet model during M2 and record it in the pilot checklist.
Engineering default; product may tighten the byte cap
D8
Assignment removed, or inspection cancelled, while this tablet has unsent work
The server still accepts that inspector's pending field and photo ops, writes them into the record and the audit, and does not reopen a cancelled inspection or give the submitter accept rights after reassignment. submit from a non-assignee is rejected. After those ops are acknowledged, the tablet purges the local copy. The new assignee sees the delivered notes. If the server rejects the payload as invalid, the tablet keeps the data, shows "sync blocked," and does not delete it.
Product, before pilot
D9
What an audit row contains
Every accepted mutation writes one audit row in the same database transaction as the version bump: actor, role, device id if any, op id if any, client timestamp if any, server timestamp, base version, new version, op type, field path if any, before value, after value. A failed audit insert rolls the mutation back. Analytics and client telemetry do not get values, names, addresses, or free text. Server retention of the audit log is unchanged by this plan. Historical edits from before the field map exists are not reconstructed.
Compliance, before real customer data
D10
Kill switch
A per-inspector and per-device flag, default off. Off means the tablet enqueues no new offline ops and will not open the local-first form for new work. If an outbox is non-empty, the app keeps the sync path on and asks the inspector to upload. The flag does not delete local work. Pilot rollback uses the flag, not a binary downgrade.
Adopting this plan
D11
Shared tablets
Pilot tablets are one inspector each. A second person signing in cannot read the first person's encrypted inspection data. Unsent work stays until the owning inspector signs back in and syncs, or until remote wipe.
Product; default stands for the pilot
User flow
Visible states on an inspection card and on the open form:
State
Meaning
Not on this tablet
No snapshot yet. Not editable.
Available offline
Snapshot and pinned schema are stored. No local edits.
Saved on this tablet
Local commit succeeded. One or more ops are not acknowledged.
Syncing
A sync pass is in progress.
Needs review
This inspection is paused on an overlapping field.
Submitted, waiting to upload
submit is queued. Photos, signature, or earlier ops are still outstanding, or the network is down.
Removing local copy
Server has accepted. Purge is running.
Removed from this tablet
Customer content is gone. The card may show an opaque id and "Submitted."
Prefetch, online. The inspector unlocks the tablet and the app. When the network is actually usable (a successful authenticated request, not merely a radio association), the app refreshes the session and prefetches every inspection assigned to that inspector: header, integer version, status, pinned schema, field map, attachment metadata. Photo bytes already on the server download until the free-space floor; metadata is still stored if bytes are skipped, and missing reference images show as placeholders. Each inspection that has a committed snapshot and its schema becomes "Available offline." Prefetch does not change server status.
Open. If no snapshot is stored, the form stays closed, including when the radio just failed mid-download. There is no empty editable form. If a snapshot exists, inputs enable only after that snapshot transaction is committed. Opening does not by itself enqueue an op.
Edit, offline or online. Field changes, photo capture, photo delete, and signature capture commit in a local transaction and show "Saved on this tablet" immediately. The 41st non-deleted photo is refused with the count visible (for example, "40 of 40"). The signature pad works offline. Replacing the signature before submit is allowed. Leaving the form, force-quitting, or rebooting does not drop committed work.
Intermittent signal. A sync pass runs when an authenticated request can get through, and again after a failure with backoff capped at 15 minutes plus jitter. The pass uploads ready blobs, pushes ready ops, then pulls. A drop mid-upload resumes the blob. A drop mid-batch retries the same op ids. The inspector can keep editing during sync.
Supervisor edit. The supervisor changes fields on the web. The server diffs the save, bumps the inspection version once, and audits the changed fields. On a later pull, fields the inspector did not change update quietly. Fields both sides changed move the inspection to "Needs review." The inspector sees, for each such field, the label, their value, and the supervisor value, and picks one or types a replacement. Other inspections keep syncing.
Submit. The inspector signs (if they have not) and confirms submit. The confirmation states that after the server accepts, this tablet will remove the inspection and further corrections happen with a supervisor. Local validation requires the pinned schema's required fields, a signature, and a photo count from 0 through 40 (and any schema minimum). Failure leaves the form open. Success queues submit and shows "Submitted, waiting to upload." While that op is still pending and unsent, the inspector may withdraw it and keep editing. After the first send, withdraw waits on the result: an accept purges, and a rejection returns the form to editing. The client sends submit only after that inspection's blobs are confirmed and its earlier ops are acknowledged.
Accept and purge. The successful submit response, or a retry that returns the original accept, switches the card to "Removing local copy" and purge runs immediately. The inspector sees "Removed from this tablet." They cannot reopen the content on the tablet.
Reassignment or cancel. If pull says the inspection is no longer theirs, or is cancelled, behavior follows D8. Clean local copies delete on that pull. Dirty copies stay visible as "sync required" until the delivery attempt finishes, then purge.
Lock and sign-out. After the offline session limit, or on sign-out, the form locks and customer content stays encrypted. Nothing is purged because of lock or sign-out.
stateDiagram-v2
[*] --> NotOnTablet
NotOnTablet --> AvailableOffline: prefetch commits snapshot
AvailableOffline --> SavedOnTablet: local edit commits
SavedOnTablet --> SavedOnTablet: further local edits
SavedOnTablet --> Syncing: usable network
Syncing --> SavedOnTablet: ops still open
Syncing --> AvailableOffline: all ops acked and still in progress
Syncing --> NeedsReview: overlapping field
NeedsReview --> SavedOnTablet: inspector resolves
SavedOnTablet --> SubmitQueued: submit confirmed
SubmitQueued --> SavedOnTablet: withdraw before send
SubmitQueued --> Removing: submit acked accepted
Removing --> Removed: purge commits
Scope boundaries
In this plan:
Local-first open, form fill, up to 40 photos, and one signature for inspections assigned to the signed-in inspector and already prefetched.
The same local-first write path when the network is up, so a drop after open cannot discard notes.
Opportunistic, resumable, idempotent sync to versioned REST through one new sync API and blob upload.
Field-level merge, with manual resolution of overlaps on the tablet.
Server audit history for mutations, including those that were captured offline.
Local purge of customer data after acceptance, and purge of a clean copy when the server reports accept, cancel, or loss of assignment.
Migration of the existing SQLite cache into a durable outbox store.
A pilot flag, kill switch, metrics, and the supervisor web server's version check and field diff. The supervisor interface stays a normal form save.
Outside this plan:
Real-time co-editing, presence, shared cursors, character-level merge, and operational transforms.
Creating, cloning, or deleting inspections on the tablet.
A supported workflow of two tablets editing one inspection.
Offline use for supervisors, or a supervisor tablet app.
Video, chat, and floor-plan markup.
A new public API style for the web. Existing REST stays.
Changing how long the server keeps audits or inspection records.
Editing an inspection on the tablet after accept.
Peer-to-peer sync between tablets.
Recovering notes the current cache already lost.
A general offline mode for the rest of the mobile app.
Local data model
SQLite remains the only local database. Its role becomes: replica of prefetched server state, working copy, and outbox. Photo and signature bytes live as encrypted files in the app sandbox, referenced by the database, excluded from the system photo library and from device backup. User-data transactions use SQLite's full synchronous mode. The database user version gates migrations.
Customer content (header, field values, photo bytes, signature bytes, signer name, address, names) is encrypted with the per-inspection key from D5. Operational columns (ids, versions, states, timestamps, error classes) are not a license to copy values into them.
device_sync
One row. device_id (stable, created on first launch, stored with the keystore material), pull_cursor, server_time_offset_ms, last_online_auth_at, schema_catalog_version.
Primary key inspection_id. server_version (integer last applied), status (assigned, in_progress, accepted, cancelled), schema_id, schema_version, assignee_id, header_encrypted, accepted_at (server time, nullable), local_phase (clean, editing, conflict, submit_queued, purge_pending, purged), purge_deadline (server accepted_at + 24 hours, nullable), purged_at, dek_id (null once the key is destroyed). Header is server-owned display data (site label and address) and is replaced on pull. Inspector edits go through schema fields.
fields
Primary key (inspection_id, field_path). server_value_encrypted, server_field_version (integer), local_value_encrypted, dirty, conflict_server_value_encrypted nullable. A dirty field still holds local_value after a pull. Purge nulls both value columns.
attachments
Primary key client_id (UUID). inspection_id, server_id nullable, kind (photo or signature), field_path nullable (a photo may illustrate one question), sha256, byte_size, local_relpath nullable, upload_state (pending, uploading, confirmed, remote_only, purged), captured_at_client, tombstone. Active photos are kind = photo and tombstone = 0. Their count is at most 40. Active signatures are at most one.
outbox
Primary key op_id (UUID). inspection_id, seq (null until first send; then a per-device monotonic integer, never reused or reassigned), op_type (upsert_field, add_photo, delete_photo, set_signature, submit), field_path nullable, attachment_client_id nullable, base_version integer, payload_encrypted, payload_sha256, state (pending, inflight, acked, conflict, rejected, superseded), attempt_count, last_error_class, created_at_client.
seq is assigned in send order when an op first enters inflight, and only for ops that are actually sent. A pending op with a null seq can be updated or removed without leaving a gap. Payload and seq are immutable once state has been inflight. A later edit to the same field while the op is still pending with a null seq updates that row in place (same op_id, new payload and hash). After the first send, a further edit is a new op. Among live ops on an inspection, submit is last in send order. Withdrawing submit deletes that row only while it is pending with a null seq, and returns local_phase to editing.
blob_jobs
Primary key client_id. bytes_confirmed, attempt_count, last_error_class. Independent of outbox order so a large photo does not sit in front of a text field at the server. The add_photo or set_signature op is eligible to send only after upload_state = confirmed.
sync_receipts
Primary key op_id. inspection_id, result (applied or rejected), result_version, server_time. No field values and no names. Written when an op reaches a terminal server result. Survives purge so support can see that work was acknowledged. Retain 45 days, then delete. This table is operational, not the regulatory audit.
Invariants
The form cannot enter an editable state unless an inspections row and its schema row are committed and status is assigned or in_progress and local_phase is not purged.
A local content change and its outbox row commit in one transaction. The file for a new photo is durably written before that transaction references it. Startup deletes unreferenced temp files.
Purge runs only when status is accepted (or D8's delivery has fully acknowledged, or the copy is clean and the server has dropped it) and no live outbox rows or unconfirmed blobs remain for that inspection.
After purged_at is set, header, field values, attachment rows' paths, signer name, and the inspection key are gone.
Client telemetry never selects encrypted columns.
Sync protocol
Delivery is at least once. Effect is exactly once per op_id. The server stores op_id and payload_sha256 in the same transaction as the mutation. A repeat with the same id and hash returns the original result and does not bump the version again. A repeat with the same id and a different hash is rejected with payload_mismatch and is not applied.
Resources the server maintains
These fit the existing integer versions. M1 maps them onto the current routes.
Inspection. Integer version. Status. Assignee. Schema id and version. Header.
Field map.field_path → {value, modified_version}. modified_version is the inspection version at which that field last changed. Initial backfill sets every field's modified_version to the inspection's current version.
Attachment metadata. Client id, server id, kind, checksum, byte size, optional field path, tombstone, inspection version at which it was added or tombstoned.
Blob store. Bytes addressed by server attachment id, checksum verified, not publicly readable.
Processed ops.op_id, payload hash, result, resulting version. Retained at least 45 days and at least until the device has purged or the inspection has been accepted for 45 days, whichever is later.
Audit log. As in D9.
All of these writes for a single op go through one domain service used by the sync API and by the existing web REST handlers. Web saves send the inspection version they displayed. The server rejects a stale web write. The web client refreshes and the supervisor retries. The server diffs a successful web save against the field map, updates only changed paths, bumps version by one, and writes audit rows for the changed paths.
Push
POST /v1/inspection-sync, authenticated.
Request body:
device_id, client_time
push: an ordered list of ops the client believes are eligible. Each op: op_id, seq, inspection_id, base_version, op_type, field_path if any, attachment_client_id if any, payload, payload_sha256, client_ts
pull: cursor (opaque, nullable), held_inspection_ids (ids currently stored on the device), known_schemas (schema_id, schema_version), limit
The client sends at most 50 ops or 256 KB of JSON, whichever comes first. A send batch is a creation-order prefix of the live ops on that inspection. As those ops enter inflight, each null seq is set from the device counter. The client does not put an op in the batch until:
every earlier live op on that inspection is already terminal or sits ahead of it in the same batch, and
for add_photo and set_signature, the blob checksum is already confirmed, and
for submit, every other live op on that inspection is acknowledged and every active attachment blob is confirmed.
The server processes push in list order, locking each inspection row. For each op:
Condition
Result
op_id already stored with the same hash
duplicate and the original result. Version unchanged.
op_id stored with a different hash
rejected / payload_mismatch.
A smaller seq from this device_id for this inspection was never stored and is not earlier in this request
blocked / sequence_gap. Later ops for that inspection in this request are not attempted.
op_type is upsert_field or delete_photo or set_signature, and the target's modified_version is greater than base_version
conflict, including server version, field path, and server value. The op id and seq are stored as a terminal conflict result so the sequence is consumed and a retry returns the same result. The field is not changed, the version is not bumped, and no before/after audit row is written.
op_type is upsert_field and modified_version ≤ base_version
Apply value, set field modified_version and inspection version to the new version, audit, store op, applied. If status is assigned, set in_progress. If status is cancelled, keep cancelled. The actor does not have to be the current assignee (D8).
add_photo and the client id is new, checksum matches a stored blob, active photo count after add ≤ 40
Apply metadata, bump version, audit, applied.
add_photo and that client id is already stored with the same hash
applied idempotent result. No second photo and no version bump.
delete_photo of a known id whose attachment modified_version ≤ base_version
Tombstone, bump version, audit, applied. A repeat tombstone is an applied no-op.
set_signature and the slot's modified_version ≤ base_version, checksum matches a stored blob
Store the signature metadata, bump version, audit, applied.
set_signature and the blob is not durable yet
blocked / attachment_not_ready. The id and seq are not consumed.
add_photo and the blob is absent or checksum differs
blocked / attachment_not_ready or rejected / checksum. Not stored as a final rejection if the blob is simply not durable yet (attachment_not_ready is retryable and does not consume the id).
add_photo would make 41 active photos
rejected / photo_limit.
delete_photo of an unknown id
applied no-op if the tombstone is already there or the id was never accepted; otherwise rejected / unknown_attachment.
submit and actor is the current assignee, status is assigned or in_progress, signature blob is durable, required fields pass, active photos ≤ 40 and ≥ schema minimum, no referenced blob missing
Set accepted, set accepted_at to server now, bump version, audit, applied.
submit and a precondition fails
rejected with a stable code (not_assignee, bad_status, signature_required, field_required, photo_limit, blob_missing). The id and seq are consumed so the same broken payload is not reapplied. The client must enqueue a new submit after fixing local state.
A conflict or a terminal rejection does not roll back earlier ops in the batch that already applied. Terminal results (applied, conflict, rejected, and the stored result behind duplicate) consume op_id and seq. Retryable blocked results, including attachment_not_ready and sequence_gap, do not. The response lists one result per attempted op. HTTP 200 carries per-op results. Authentication runs before any apply. HTTP 401 means the token is dead and nothing in that request was applied. A 5xx means the client retries the same ids and the same seqs; anything not committed was not stored.
base_version is the inspection version the edit was based on, not a per-field guess by the client. The server compares that number to the field'smodified_version (or the attachment's modified version). An inspection version that moved because some other field changed is not a conflict.
The same status rule applies to an applied add_photo, delete_photo, and set_signature: assigned becomes in_progress, and cancelled stays cancelled.
Pull, in the same response
server_time
next_cursor
changed: inspections newly assigned or updated since the cursor and still relevant to this user. Each includes id, version, status, assignee, schema id and version, accepted_at if set, header, full field map (path, value, modified version), attachment metadata. No blob bytes.
dropped: held ids that are no longer assigned to this user, or are cancelled or accepted, if they were not already fully described in changed. Reason code included.
schemas: bodies the client asked for by omission from known_schemas and that the changed inspections need.
The client stores next_cursor only in the same local transaction that finishes applying the page. Replaying a page is safe: a field snapshot at a given version applied twice matches.
Pull application while local work exists:
Not-dirty field: replace local and server shadows with the pulled value.
Dirty field whose pulled modified_version ≤ the pending op's base_version: keep the local value. Update the inspection's server_version only as a high-water mark. Do not rewrite the op's base_version.
Dirty field whose pulled modified_version > the pending op's base_version: keep the local value, store the server value on the conflict column, set local_phase = conflict, stop pushing further ops for this inspection until resolution.
Do not clear a dirty value because the inspection version increased.
Client send order in a pass: confirm blobs that are ready (limited parallelism, two at a time), push eligible ops, then apply the pull section of the response. If a push result is conflict, apply it before using the pull to fill the review screen.
Backoff uses last_error_class. sequence_gap, attachment_not_ready, and network failures are retryable and keep the assigned seq. payload_mismatch pauses further pushes for that inspection until the op is superseded by a new op id; it must not be skipped. A validation rejected on submit returns the form to editing. Other terminal rejections surface "sync blocked" on that inspection.
Blob upload
POST /v1/inspection-attachments with client_id, inspection_id, kind, sha256, byte_size, content_type. Response: server_id and a short-lived upload instruction bound to that id. Repeating the same client_id and hash returns the same server_id.
Resumable byte upload. The server records the contiguous confirmed prefix.
POST /v1/inspection-attachments/{server_id}/complete with the hash. The server checks the stored bytes. Mismatch deletes the partial and returns checksum. Match sets the blob durable.
Upload instructions expire; the client requests them again and resumes. Bytes are not fetched by unauthenticated URLs.
Resolution rewrite
When the inspector finishes the review screen, the client, in one transaction:
For each conflicted field, if they kept the server value, drops the local dirty value and supersedes the unapplied op.
If they kept theirs or typed a new value, supersedes the unapplied op and inserts a new upsert_field whose base_version is the pulled inspection version and whose payload is the chosen value.
Leaves unapplied ops for fields the server did not change, still with their original base_version.
Clears local_phase to editing once no conflict columns remain.
A superseded op_id is not sent again. If the server already stored it as conflict, that terminal result stays, and the replacement is a new op_id with a new seq assigned when it is sent. Applied ops stay applied. This is a rebase of unacknowledged work.
Time
Each successful sync stores offset = server_time - device_clock. Purge deadlines use accepted_at from the server. Audit order uses server_time. Client timestamps are stored beside it. The server still accepts an op whose client_ts is implausible; it records the claim and sets an implausible_client_time flag on the audit row. It does not reject the inspection for clock skew.
Conflict policy
One writer is supported on the tablet, plus supervisors on the web. That is the overlap the product has to survive without becoming a collaborative editor.
Rules:
The unit of conflict is one field path, one signature slot, or one photo identity.
Text is not merged inside the string. If both sides changed the same field, the inspector picks the whole value.
Photo adds with distinct client_ids all stay. Deletes are tombstones. Two deletes of the same id are one delete. A delete conflicts only if that attachment's metadata modified version is newer than base_version and the server state is not already the tombstone the client is trying to write.
The signature is one slot. The same rule as a field applies. The later resolved set_signature is the signature that submit requires.
The photo cap is enforced at add time and again at submit. A merge that would exceed 40 rejects the op that crosses the cap; the inspector deletes down and retries with a new op.
Conflict pauses push for that inspection only.
The web client does not grow a merge UI in this plan. A stale supervisor save fails the version check; the supervisor reloads the form, which by then includes any ops already applied, and saves again.
Unknown schema types: the inspection is readable where the app understands it, and submit stays disabled with "Update the app to submit this inspection." The client does not drop unknown fields on sync.
Worked example. Tablet pulled version 3. Supervisor sets field A; server version is 4; field A modified_version is 4; field B remains at 2. Tablet, still based on 3, sets field B and field A.
Field B applies, because 2 ≤ 3. Inspection version becomes 5. Both the supervisor's A and the inspector's B are stored.
Field A returns conflict. The tablet shows the supervisor's A and the inspector's A. Nothing in that op was written.
Attachments
Captures go to the sandbox through the encrypted file path. They are not written to the shared photo library. Share-sheet and screenshot export of an inspection are disabled where the platform allows, and the pilot device profile turns them off.
A photo row is created only after the file is durable. The outbox add_photo follows blob confirmation, so the server does not audit a photo it cannot read.
Deleting a photo whose add_photo op is still pending with a null seq removes the local file, the blob job, and that op. No server op. Once the op is in flight, the tablet waits for the result and then sends delete_photo if the add was applied.
Deleting a photo already applied enqueues delete_photo.
The signature uses the same blob path with kind = signature and is not counted toward 40. Photos that originated on the web count toward the same cap of 40.
Checksums are SHA-256 computed on device before upload. The server recomputes.
Interrupted uploads resume from bytes_confirmed.
Reference images that originated on the server are prefetched best-effort. Failure to fetch a reference image does not block form fill or submit.
Disk pressure: stop new captures under the free-space floor, keep the outbox, and show why. Do not delete queued customer data to free space.
submit is the completeness gate. The server accepts only when every referenced blob is durable. Until then the tablet retains the local copy and shows "Submitted, waiting to upload."
Security
Pilot enrollment requires the MDM baseline in A8: device encryption, screen lock, remote wipe, backup of the app data store disabled, OS version on a floor the app already supports.
The app holds session tokens in the platform keystore. The offline window follows D4. After 12 hours without online authentication the UI locks. Data and outbox remain. Sync that receives 401 locks the sync action, keeps the outbox, and asks for online sign-in.
Each inspection's field values, header, and files use a distinct data key wrapped by a hardware-backed device key. Purge deletes the wrapped key first, then rows and files, inside the purge transaction's cleanup. Flash wear means overwrite is not the control; destroying the key is.
SQLite, its journal and WAL, temp capture files, and thumbnails are covered by the same key and the same purge file list. Startup removes temp files with no attachment row.
TLS for every sync and upload call. This plan does not add certificate pinning; a pinned key plus a wrong device clock is a way to trap a tablet that cannot sync, which fights the deletion requirement. If the app already pins, keep the existing policy and test clock skew explicitly.
Upload URLs are authenticated and short-lived. Blob reads for prefetch use the same session.
The server rejects a mutation it cannot audit (D9). There is no admin path that updates an inspection field without the domain service.
device_id is an audit and support attribute, not a proof of identity. Identity is the user token. Ops name the user from the token, and the device id from the body, and the server stores both.
A jailbroken or rooted device is recorded as a signal on sync and excluded from the pilot cohort. It is not the only control.
Telemetry, crash logs, and support exports omit header text, field values, signer names, addresses, photo bytes, and signature bytes. They may include inspection id, op id, versions, op type, error class, byte counts, and device id.
Lost tablet: MDM remote wipe remains the control for a device that will not sync. Local data is unreadable without the device key and the lock credential. That is the residual case named under A1.
The server audit log holds customer values because D9 says reconstruction is the point of the audit. Access to audit rows follows the same staff authorization as the inspection record. This plan does not copy that log onto the tablet.
Migration and backfill
Tablet. On first launch of the build that contains this store:
Read the SQLite user version.
Create the new tables beside the old cache.
Do not copy cache rows into fields, attachments, or outbox. A4 says they are not unique work, and promoting them would manufacture false dirty state.
Drop the cache tables once the new schema is committed, or leave them unread if a later binary still needs to open the file during rollback testing. They are not a source of truth either way.
Set a local "prefetch required" bit. Until one successful prefetch completes, every assigned inspection renders as "Not on this tablet," including if the radio is down. The inspector sees that offline work is unavailable until the tablet has synced once on this version.
If the upgrade transaction fails, the app stays on the online read-through behavior and reports the migration error. It does not half-open the new form.
Server, before any pilot device can enqueue an op.
Ensure every inspection row has an integer version, backfilling missing values to 1.
Build the field map from the current stored form. Set each modified_version to that inspection version. Write one audit event field_map_baseline with actor system and no per-field before/after dump (the current row remains the current value; this plan does not invent a history the old system did not store).
Add processed-op storage, attachment client ids and checksums, and the audit columns in D9.
Route existing web writes through the domain service. Confirm there is no second writer that updates inspection columns directly.
These server changes are safe while the client flag is off: behavior for a web save that does not touch fields is a version bump only when the diff is non-empty. A stale write newly returns a version conflict instead of overwriting. That is intended.
Rollback of the migration. A forward migration that discarded the cache can be followed by an app downgrade only when every pilot device outbox is empty. The downgraded app may recreate a cache by refetching. It must not be pointed at the new tables and interpret them as the old cache.
Observability
No metric label or log property may contain customer content. Allowed dimensions: app version, cohort, op type, error class, HTTP class, platform.
Client:
Sync pass result and error class.
Outbox depth and age of the oldest live op.
Blob failures, retries, and bytes confirmed.
Conflict count, by inspection, without field path in the metric (field path may sit in access-controlled server logs).
Purge result: purged, overdue, blocked_unsent.
Prefetch: inspections stored, inspections skipped for space, schema missing.
Photo-count histogram at submit time (counts only).
Offline session length.
Server:
Op apply latency, duplicate rate, conflict rate, rejection codes.
Web version-conflict rate (stale If-Match or equivalent).
Audit transaction rollback count. Page on any non-zero rate; a mutation that cannot audit is a failed mutation, and a rising count means inspectors cannot save.
Blob complete failures and checksum failures.
Time from accepted_at to device-reported purge, for devices that have reported.
Alerts for the pilot:
Audit rollback count > 0 in a five-minute window.
A device that has pulled accepted_at and has not reported purge, and whose accepted_at is older than 24 hours.
Outbox age over 8 hours while the device is completing other authenticated requests (stuck queue, as opposed to a basement).
Sync error rate excluding conflict and user-caused rejected above the dogfood baseline by a sustained margin. Set the numeric page threshold from dogfood rates before the pilot cohort expands; do not invent the threshold here.
Support can read, for an inspection id: op ids, states, versions, error classes, accept time, purge time. They cannot read field values from telemetry. Server inspection and audit screens remain the place to see content, under existing staff access.
The pilot dashboard shows cohort size, inspections submitted through the local-first path, conflict rate, purge lag, stuck outbox count, and open data-loss incidents. It does not show customer fields.
Staged rollout
Stages are ordered. Exit criteria are the gate. There is no date on any stage.
Stage 0 — server dark. Ship field map, domain service, audit transaction, blob API, and sync API to production with the client flag off for everyone. Web saves go through the domain service. Exit: web saves bump version and write field-level audit; a deliberately stale web write is rejected; processed-op replay in a production-like environment does not double-apply; no client has the flag.
Stage 1 — synthetic dogfood. Flag on for staff accounts and non-customer inspections only, until D1 and D9 are confirmed. Script the basement case: prefetch, radio off, edit, photos, signature, kill app, reboot, radio on, accept, purge. Exit: the script passes on the pilot hardware; migration from a cache build does not present cache rows as dirty; kill switch drains a prepared outbox and then blocks new local-first opens; telemetry sample contains no field values.
Stage 2 — pilot cohort. Product names the inspectors and devices. Real assignments. Support has the state table in this plan. Exit, all of them:
No open incident where a committed local note was lost because of connectivity.
Every accepted inspection whose tablet completed a sync within 24 hours of accepted_at shows a purge reported inside that window.
Any accepted inspection that missed the window has a written compliance disposition before the cohort grows.
Conflicts that occurred were resolved on the tablet, or are listed as product issues with a decision to continue or stop.
Photo rejected / checksum / permanently failed upload count is understood. A permanent failure rate above 2% of photos in the cohort stops expansion until the cause is fixed. (This threshold is a pilot gate, not a capacity plan.)
Audit rollback alert has stayed at zero.
The kill switch has been exercised once in dogfood and the runbook matches what the device did.
Stage 3 — broad company rollout. Flag default on for devices that meet the MDM bar. The kill switch stays. Exit: stage 2 criteria hold for the first broad slice the same way, and the online read-through write path is no longer reachable from the inspection form.
Stage 4 — cleanup. Remove unread cache code and any duplicate write path. Sync API and web REST still share the domain service.
Stop the rollout and return to the flag-off posture (draining existing outboxes) if audit rollbacks are non-zero, if a purge gap has no compliance disposition, or if a confirmed data-loss bug is open.
Rollback
Flag off is the rollback. Tablets with an empty outbox behave as online read-through again and do not open a half-cached form as editable. Tablets with a non-empty outbox keep uploading until it is empty or until an op is rejected or conflict, which support handles from the state table. The flag never deletes the outbox.
Server endpoints stay up while any client might still push. Disabling the route while outboxes are dirty strands customer data on tablets and fights D1. The dark behavior is "no new clients enqueue," which is the flag, not an API outage.
Do not ship a binary downgrade to a device with live outbox rows or unconfirmed blobs. The old binary does not know how to upload them. If a binary rollback is ever required, drain first, then downgrade, then let the old app refetch a cache.
Bad server apply. Because every apply stored the op payload hash and the audit before/after, an inspection's field map can be rebuilt to a prior version by replaying audit rows forward from field_map_baseline or by reversing a single audited op under a new compensating op. Compensating ops are a support tool using the domain service, so they are themselves audited. They are not a tablet feature.
Purge bug that deletes before ack. Treat as data loss. Stop the cohort via the flag. Do not tell inspectors to keep working offline until the purge guard is fixed and tested.
Testing
Tests are part of the milestone that introduces the behavior. A milestone is not done while its rows fail.
Case
Pass
Milestone
Prefetch, enable radio off, edit three fields, kill app, reboot, radio still off
Values present and marked saved on tablet
M2
Radio dies after open, before any explicit save gesture
Edits made after the snapshot commit are present
M2
Upgrade from a cache-populated build
New store empty of dirty fields; inspections not editable until prefetch; cache text does not appear as a local edit
M2
41st photo
Refused; first 40 remain
M2
Submit with a required field empty or signature missing
Submit not queued; form remains editable
M2
Queue submit, withdraw before the op is sent
Submit op gone; editing restored. Withdraw is unavailable while the op is in flight
M2
Blob upload interrupted at a known byte offset
Resumes; one blob; checksum matches
M2
Same op_id and hash delivered twice
One version bump; second result duplicate
M1
Same op_id, different hash
No apply; payload_mismatch
M1
Supervisor changes field A, offline inspector changes field B, then sync
Both values stored; one new version for B on top of A
M1, M3
Both change field A
A's op is conflict; B's non-overlapping op can still apply; tablet shows A for resolution; chosen value syncs as a new op
M1, M3
Property check: random field sets from both sides
Every non-overlapping field survives from both writers; an overlapping field is never stored from both without a resolution op
M1
Web save with a stale version
Rejected; reload then save applies; no direct SQL writer bypasses this
M1
Audit insert forced to fail
Field unchanged, version unchanged, op not stored
M1
submit with a missing blob, then complete the blob and retry
First result does not accept; after the blob is durable, accept happens once
M1, M4
Lost HTTP response after a committed submit
Retry returns duplicate and the original accepted_at; tablet purges once
M4
Accept response applied locally
Header, fields, files, thumbnail, and data key unreadable; sync_receipts has no values
M4
Device clock set backward across purge
Purge still happens from server accepted_at; deadline does not move with the device clock
M4
Reassignment with a dirty field
Field is audited on the server; assignee can see it; old tablet purges after ack; old tablet's submit is not_assignee
M1, M3
Cancel with a dirty field
Content op audited; status stays cancelled; tablet purges after ack
M1
Token expired offline
Edits within the 12-hour window still commit locally; sync does not start until online auth; data still present
M2
Sign-out
Next user cannot read the inspection; previous outbox still there after the owner signs back in
M2, M5
Flag turned off with a non-empty outbox
No new ops; existing ops upload; then the form is online-only
M5
Telemetry and crash payload from a sync with text, a photo, and a signature
Payload has ids, versions, counts, and error classes only
M5
Disk full
Capture refused; existing outbox intact
M2
Unknown future field type marked required
Submit disabled; known fields still sync
M3
Schema version pinned
A newer published schema does not change an in-progress inspection's local form
M2
Also run the merge and purge tests under a proxy that drops connections after the server has committed and before the response, and again mid-blob.
Milestones
Dependencies are the schedule. M1 and the local capture portion of M2 can proceed in parallel once M0 has recorded the engineering defaults. Real customer data waits on compliance confirmation of D1 and D9.
Milestone
Depends on
Work
Acceptance
M0 — decision record
None
Record adopt or override for D1–D11. Name the compliance owner for D1 and D9. Inventory current REST resources, version columns, and write paths (the M1 discovery).
The decision table has a status for every row. Any override is written into this plan before code that would contradict it. Gaps found in the API inventory are listed as M1 tasks.
M1 — server domain path
M0
Field map and baseline backfill. Domain service. Web writes use it and honor integer versions. Audit transaction. Processed ops. Blob initiate, resume, complete. POST /v1/inspection-sync as specified. Flag still off for customer devices.
M1 rows in the test table pass. A replayed op does not double-apply. A non-overlapping pair merges. An overlapping pair conflicts. A forced audit failure leaves the record unchanged. A stale web write is rejected. No unversioned writer remains.
M2 — tablet local-first store
M0 for capture and schema; frozen M1 request shape before the sync client is finished
New SQLite store, migration that discards cache-as-truth, encrypted files, outbox, coalescing of unsent field ops, photo and signature capture, caps, prefetch, sync worker, submit and withdraw. Flag off in production.
M2 rows in the test table pass on a tablet build. The disconnect-after-open case retains notes. An inspection with no snapshot is not editable.
M3 — conflicts and assignment changes
M1, M2
Review screen. Pull application rules. D8 delivery. Submit lock while a conflict is open.
M3 rows pass. An inspector can resolve an overlap without a support tool. A second inspection on the same tablet keeps syncing while the first is in review.
M4 — purge
M1, M2, and a recorded D1
Immediate purge once an accept result is stored locally. Startup finishes an interrupted purge. A lost accept response is recovered by retrying submit, then purging from that stored result. Crypto-erase, backup exclusion, purge metrics.
M4 rows pass. A test that plants an accepted inspection with an unacknowledged op does not purge. A lost response purges only after the duplicate result is stored. Customer bytes are unreadable after the successful path.
M5 — operate it
M1–M4
Metrics, pilot dashboard, alerts, kill switch behavior, support notes that use the state table in this plan.
M5 rows pass. A sampled telemetry payload has no customer content. Flag-off drain matches the rollback section. The audit alert fires in a forced failure test.
M6 — dogfood then pilot
M5; D1 and D9 confirmed before real customer data
Stage 1, then stage 2, using the exit criteria in staged rollout.
Each stage's exit criteria are recorded as met or the cohort is stopped. No calendar commitment.
M7 — broad rollout and path removal
M6 exit
Stage 3 and stage 4.
The inspection form has a single write path: local commit, then sync. The read-through write path is gone. The kill switch still drains and disables.
M0 does not end by silence. Unconfirmed D1 or D9 blocks M6's real-customer stage and does not block synthetic M1–M5.
Residual risks the milestones do not remove
A tablet that commits submit on the server, loses the response, and never powers on again still holds encrypted customer data past 24 hours until someone wipes it. A1 states this residual explicitly.
A long period with no usable uplink delays acceptance itself, so the local copy remains by design (A3). The dashboard's outbox age is how that is seen.
Two tablets for one inspector are unsupported. The second device produces ordinary conflicts, which are easy to misread as data loss if the pilot ignores A7.
Server audit rows contain customer values for as long as the server retention schedule says. That is D9, and it is a different store from the tablet.
Compensating a bad apply repairs server state. It cannot repair a tablet that already purged, because the tablet no longer has a copy. Post-accept corrections stay on the server.
grok-4.7/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 8.25 / 10graded blind as submission E
A compact, correct plan. Acceptance waits for a purge receipt or an enforceable lease, and the plan says plainly that an app-only background job cannot guarantee deletion. Conflicts use a strict record-level check, and the supervisor resolves them through a web conflict package, which may produce a lot of review work and leaves the inspector blocked. The milestone table is sound but brief.
Strengths
Purge-before-accept, with a lease alternative and an honest statement of platform limits
Clear lifecycle states that separate saved, synced and accepted
Rollback keeps sync and purge paths alive and blocks managed downgrade until outboxes drain
A decision log that records the alternatives considered
Weaknesses
A thinner local data model and fewer details on clock skew
Whole-record conflicts go to supervisors for every overlap, an operational burden that is only partly acknowledged
Milestone acceptance criteria are summary-level
Evidence the grader checked
Recommended design paragraph on purge receipt or lease
Sync protocol step 6 'No automatic merge in the pilot'
Milestones table M0–M7
Objective checks
Files
gpt-6-luna/06-product-planning/RESPONSE.md
Three riskiest assumptions
The 24-hour deletion rule can be enforced on a managed tablet that is offline or powered off. The brief requires local customer data removal within 24 hours after acceptance, but does not describe the tablet OS, background execution guarantees, MDM wipe/lease controls, or whether web acceptance can be delayed. If deletion depends only on the app running, this assumption may fail. The plan recommends making server acceptance wait for a purge receipt or enforceable offline-data lease expiry. Validate this behavior on the actual managed devices before the pilot; if it cannot be guaranteed, product/compliance must change the acceptance flow or clarify the requirement.
Integer record versions are sufficient to prevent silent loss while inspectors and supervisors edit the same inspection. The brief confirms versioned REST APIs but not conditional-write semantics, transaction boundaries, idempotency, or any audit of web writes. The plan assumes an atomic compare-and-update, rejects stale writes, and routes the preserved local candidate to a supervisor conflict review. Confirm that all web/API write paths can participate and name the conflict resolver before implementation. Otherwise drafts could be overwritten, duplicated, or stranded.
The server can durably retain the full inspection, attachments, signature, and required audit history before the tablet purges its copy. The brief requires an audit history and up to 40 photos, but does not define audit contents/retention, signature evidence requirements, upload limits, or attachment storage behavior. The plan assumes resumable, hash-verified uploads and an append-only server audit acknowledged before local deletion. Compliance/product and backend owners must agree on retention, payload limits, and signature representation before the API contract is frozen.
gpt-6-luna/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
gpt-6-luna/06-product-planning/PLAN.md
Offline inspection sync plan
Status: Proposed plan for implementation planning Product input:BRIEF.md only. Details not stated there are marked as assumptions or decisions. Goal: Let an inspector complete an assigned building inspection when disconnected, retain the work through intermittent connectivity, and sync it without silently overwriting a supervisor's changes.
Understanding summary
Inspectors use company-managed tablets in basements with unreliable or absent signal.
They need to open assigned inspections, complete forms, take up to 40 photos, and capture a signature offline.
Work must survive connection loss and app restarts; connectivity may return intermittently.
Supervisors may edit the same inspection through the web app.
Existing server APIs are REST-based and use integer record versions; the mobile app already has SQLite but currently uses it as a read-through cache.
An inspection's local customer data must be removed within 24 hours after it is accepted, and the server must retain an audit history.
The pilot should add dependable offline work without turning the product into a full collaborative editor.
Recommended design
Promote SQLite from a read-through cache to the durable local workspace for the inspector's assigned draft. Save each form change and an ordered outbox event in one local transaction before showing it as saved. Keep the server's integer version as the optimistic concurrency token. Sync queued draft changes and attachments idempotently when connectivity permits. If the server version has advanced, preserve both versions and stop automatic writes; send a conflict package for supervisor review rather than merging or overwriting data silently.
Treat server acceptance as a final server-side state transition. The tablet must first upload and receive durable acknowledgments for the inspection, audit events, signature, and attachments, then delete its customer data and per-inspection key and send a purge receipt. Only then should the server finalize acceptance. For a tablet that is offline when a supervisor attempts acceptance, acceptance must wait for a purge receipt or an enforceable offline-data lease to expire. This is the recommended way to make the 24-hour rule meaningful; confirm the platform and MDM can support it before pilot.
Assumptions and constraints
These are planning assumptions, not additional facts from the brief:
Offline creation is limited to assigned inspections already downloaded to that tablet while online. New assignments and unassigned customer records are not available offline.
One inspector owns a local working copy at a time. A user switch, sign-out, or reassignment must first sync or safely retain the current draft and must not expose it to another user.
A server-side accepted state is the compliance clock's start. A local “ready to submit” action while offline is not yet acceptance.
The server can add endpoints and tables compatibly, enforce the existing integer version in a transaction, and retain append-only audit events.
The managed tablet platform can encrypt local database/files, exclude them from consumer/cloud backups, and support a reliable local purge or a managed-device recovery path. This needs verification; an app-only background job cannot guarantee execution while a tablet is powered off.
“40 photos” is a maximum per inspection, with the signature stored as a separate artifact. File size, supported image formats, and compression quality remain to be decided.
Audit retention duration, audit contents, signature evidentiary requirements, and conflict-resolution ownership need approval from the relevant product/compliance owners before implementation.
User flow and states
Prepare: While connected, the inspector opens an assigned inspection. The app downloads its current form, assignment, integer server_version, and required reference data. It writes an encrypted local baseline and working copy before making the inspection available offline.
Work offline: Form edits autosave to SQLite transactionally. Each edit is also placed in a durable, ordered outbox. The UI shows local save status separately from server sync status. Photos and the signature are saved in encrypted app storage and referenced by local metadata. The app enforces the 40-photo limit.
Reconnect: Connectivity is only a hint to try sync. The app resumes queued changes and interrupted attachment uploads with exponential backoff and jitter. A successful network request is acknowledged by its idempotency key; retries do not create duplicate edits, audit events, or attachments.
Version conflict: If the server version differs from the draft's base version, the app pauses automatic writes and acceptance. It preserves the local draft and attachments, fetches the current server version, and creates an isolated conflict package. A supervisor reviews the local candidate and canonical server record in a minimal web conflict workflow, then records a resolution. This is exception handling, not live co-editing.
Submit and accept: After required fields, photos, and signature are present, the inspector submits. The server validates assignment and version, persists the final data and audit history, and confirms all artifact uploads. It returns a ready_to_accept receipt. The client removes the inspection's local data and key, then sends an idempotent purge receipt. The server marks the inspection accepted only after that receipt, or after an approved, enforceable device-lease expiry path. The app retains only a non-customer-data tombstone/receipt.
Cleanup: If acceptance is initiated from the web while a tablet still holds an offline copy, the server does not finalize acceptance until that tablet reports purge or its enforceable lease expires. The client also runs cleanup on launch and scheduled background opportunities. Cleanup removes database rows, journal/outbox payloads, photos, thumbnails, signatures, upload staging files, SQLite WAL/journal files, and encryption keys.
Suggested states are assigned_cached → in_progress_local → queued → syncing → conflict_review (when stale) or ready_to_accept → purged_pending_receipt → accepted. The state shown to the user must distinguish “saved on this tablet,” “synced,” and “accepted.”
Scope boundaries
In scope: assigned inspections; offline form entry; up to 40 photos; offline signature capture; local durable autosave; queued/idempotent sync; optimistic version checks; a bounded supervisor conflict-review/resolution path; server audit events; secure local storage and purge; rollout and operational telemetry.
Out of scope for this pilot: offline access to every customer or inspection; creation of assignments offline; simultaneous live editing, chat, presence, comments, or general-purpose field-level merge; unlimited media; synchronization across multiple inspector devices; a broad redesign of the supervisor web experience. The minimal web conflict workflow exists only to resolve a stale-version submission safely.
Local data model
Use additive, transactional SQLite schema migrations. Store customer data only for active assigned inspections.
Entity
Key fields and purpose
inspection_workspace
inspection_id, assignment_id, inspector_id, base_server_version, encrypted base_snapshot, encrypted working_snapshot, state, local_updated_at, accepted_at/purge_due_at when known. The baseline enables stale-version review; the working snapshot is the durable offline draft.
Opaque attachment ID, inspection ID, kind (photo or signature), encrypted local file reference, byte count, content hash, capture time, ordering, upload session/state. Derivatives and temporary chunks are also customer data.
cleanup_tombstone
Opaque inspection identifier, purge receipt ID, server acceptance time, and cleanup outcome only. It must not contain form values, customer identity, photos, signature data, or queued payloads.
The server remains authoritative for canonical accepted inspections and audit history. SQLite is authoritative for the not-yet-accepted local draft on that device. Never treat a stale cache row as a valid offline assignment without server-issued assignment/version metadata.
Sync protocol and conflict policy
Bootstrap: Fetch an assigned inspection snapshot plus integer version and assignment authorization. Store baseline and working copy atomically. The server must reject changes from revoked or unauthorized assignments.
Persist before acknowledging locally: On each edit, update the working snapshot and append its outbox event in one SQLite transaction. A local “saved” indicator appears only after commit. A crash or force-quit must leave either both records or neither.
Send in order: Submit events per inspection with a stable idempotency key and base_version/If-Match integer version. Server validation, canonical update, version increment, and audit append happen atomically. Return the new version and event acknowledgment. Retry of an acknowledged key returns the original result.
Upload artifacts: Create an upload session, send resumable chunks (or a resumable equivalent supported by the chosen REST service), validate final byte count and SHA-256 digest, then attach the server artifact ID to the inspection. Do not report final submission complete while any artifact is incomplete.
Stale version: Return a typed conflict response with current version. Do not apply any part of a stale write. The app pauses later events for that inspection and preserves its local state. A separate conflict endpoint stores the local candidate against its base and observed server versions without changing the canonical inspection. The minimal supervisor workflow compares the two versions and commits a deliberate resolution as a new version with its own audit event. The inspector can then continue from that resolved version if work remains.
No automatic merge in the pilot: Whole-inspection version mismatch requires review, even if the changes may be to different fields. This makes behavior predictable for regulated forms and avoids assuming field-level merge rules that the brief does not define. No signature is silently replaced or merged.
Acceptance: Server verifies required form state, the signature, all artifact hashes, authorization, and audit persistence. It issues ready_to_accept; the client purges local customer data and keys and posts an idempotent purge receipt. Server transitions to accepted and records the acceptance time only under the agreed purge/lease policy.
Attachments and signature
Enforce no more than 40 photo artifacts per inspection; count pending, uploaded, and retried photos consistently so retries cannot bypass the limit.
Persist each capture to encrypted local storage before showing it as attached. Keep stable attachment IDs across retries. Show upload state per item and allow a failed transfer to resume without taking the photo again.
Use resumable uploads with checksum verification. The API must define allowed formats, maximum bytes per file and inspection, image-processing rules, chunk/session expiry, and server retention of abandoned uploads before implementation.
Treat signatures as separate protected evidence. Preserve the captured artifact and capture metadata; do not edit or merge it during conflict handling. Confirm signer attribution, consent wording, timestamp requirements, and whether the server needs the original stroke data, a rendered image, or both.
Delete originals, previews, transformed copies, chunks, and server-side abandoned conflict artifacts according to the approved retention policy. Never put image or signature content in logs.
Security, audit, and deletion
Encrypt SQLite and all inspection files at rest with per-inspection data keys protected by the managed OS keystore. Require the tablet's normal device authentication to access app data. Exclude the database, files, and keys from unmanaged backups. Use TLS and server-side authorization for every sync, upload, conflict, and acceptance request.
Bind queued work to the signed-in inspector, device, assignment, and stable event IDs. Revalidate authorization on receipt; do not trust client timestamps or client-provided identity by themselves.
The server audit is append-only and records actor/device, event ID, affected inspection/version, action, server receipt time, client event time, artifact hashes/references, and the approved before/after detail. Audit entries must be committed atomically with canonical updates and be deduplicated on retry. The exact values to retain and retention period are compliance decisions; keep audit data out of general telemetry.
Keep logs and metrics free of customer names, field values, photos, signatures, access tokens, and raw payloads. Use opaque IDs or pseudonyms with restricted access.
Prefer a purge receipt before server acceptance. If the existing business meaning requires acceptance to occur earlier, acceptance may precede deletion by at most 24 hours only if a verified platform/MDM mechanism can ensure deletion on tablets that remain offline or powered off. Do not claim compliance based solely on a best-effort app background task. Track acceptance-to-purge age and alert before the limit; escalate any breach.
Migration and backfill
Add server capabilities additively: idempotency, conditional version writes, audit persistence, attachment sessions, conflict packages/resolution, and acceptance/purge state. Keep existing REST callers functional during the pilot; do not reinterpret the integer version.
Add a forward-only, transactional SQLite migration that creates workspace, outbox, attachment, key, and tombstone structures. Test upgrade from the current read-through-cache schema and recovery from interruption.
On first use after upgrade, refetch active assigned inspections and their server versions; do not promote arbitrary old cache entries to offline drafts. Since the current app does not support offline work, no local offline draft history needs backfill. Remove/refetch stale cache data and immediately purge already-accepted cached customer records during upgrade.
No historical server inspection backfill is proposed. Existing records continue at their current versions; new audit events begin with the new audited API path. Confirm how web/API writes enter the same audit stream before enabling the pilot.
Deploy the additive API before the mobile feature flag. Keep old-client behavior supported until the pilot cohort is upgraded. Do not ship a downgrade path that can discard an outbox or expose a new local schema to an old app.
Observability and operations
Capture operational metadata only: sync attempts/success/failure, queue age and depth, oldest pending event, retry count, attachment bytes/progress/failure, conflict count and age, server-version rejection, duplicate idempotency hits, audit-write failures, acceptance state transitions, and acceptance-to-purge duration. Use pseudonymous device/inspection IDs and redact payloads.
Provide dashboards split by app version and pilot cohort. Alert on growing queue age, repeated upload failures, unresolved conflicts beyond the agreed service window, missing audit writes, orphaned attachment sessions, acceptance purge approaching the 24-hour limit, and any cleanup breach. Support diagnostics should expose a sync state and opaque correlation ID, not customer content. Establish retention and access controls for telemetry separately from the regulated audit store.
Staged rollout and rollback
Contract and readiness: Resolve the decisions below, validate the platform/MDM purge mechanism, establish measurable pilot gates, and document a support/runbook path. No feature flag is enabled.
Server dark launch: Deploy additive endpoints/schema, idempotency, audit, attachment lifecycle, and metrics without changing existing web behavior. Verify web and mobile writes are audited and version checks are atomic.
Internal device cohort: Enable the new app path for managed internal tablets and test disconnected work, reconnect, conflict review, acceptance, and purge. Use synthetic or approved test records.
Limited pilot: Enable an explicitly enrolled, limited cohort of inspectors/assignments. Review sync reliability, photo completion, conflict burden, audit completeness, cleanup receipts, and support feedback before expanding.
Cohort expansion: Expand only after product/compliance gates are met at each step. Define numeric success thresholds from baseline and pilot data before the pilot begins; do not infer them here.
Broad availability: Enable by controlled feature flag, retaining cohort and version-level monitoring and an owner/runbook for unresolved queued work.
Rollback: Turn off enrollment and new offline-workspace creation first. Keep sync, conflict resolution, and purge paths live for every device with queued work. Do not revert server audit/version data or delete queued drafts. If submission is unsafe, pause final acceptance while preserving encrypted local work and clearly surface “not accepted”; continue purge processing for already accepted inspections. Keep API changes additive and block managed app downgrade until outboxes are drained or migrated. Resume with a corrected client/server release and reconcile all pending queues before re-enabling enrollment.
Milestones, dependencies, and acceptance criteria
Milestone
Depends on
Acceptance criteria
M0 — Product/compliance/API decisions
None
Acceptance and purge semantics, audit content/retention, conflict owner/flow, artifact limits, security/MDM capability, and pilot gates are documented and approved.
M1 — Server contract and persistence
M0
Version-conditional writes are atomic; idempotency retries deduplicate; audit entries are durable with each update; resumable attachments are hash-checked; stale writes return a typed conflict without changing canonical data; acceptance/purge states are queryable.
M2 — Local workspace and migration
M0; API contract from M1
Upgrade from the existing cache schema is transactional; active assignments can be opened offline; every local edit survives app termination; accepted cache rows are purged; unassigned/stale cache rows cannot be opened as valid drafts.
M3 — Form sync and conflict path
M1, M2
Offline edits sync in order; retries do not duplicate updates/audit entries; server version advances only on accepted writes; stale writes preserve the draft and create a supervisor-reviewable conflict without overwriting the canonical record.
M4 — Photos and signature
M2, M3, attachment API in M1
The app supports 40 photos plus the separate signature artifact; interrupted upload resumes; digest mismatch or missing artifact prevents final submission; signature is preserved across retry and conflict.
M5 — Audit, purge, and telemetry
M1–M4
Server audit is complete for mobile/web changes and conflict resolution; local purge removes database/WAL/outbox/media/key copies; acceptance is gated by the agreed purge policy; purge receipts and deadline breaches are observable without logging customer content.
M6 — Pilot readiness
M3–M5
Migration, security, failure-mode, regression, and end-to-end scenarios pass; support runbook and rollback are rehearsed; numeric go/no-go thresholds and escalation owner are set.
M7 — Pilot and expansion
M6
Pilot gate review shows no lost acknowledged edits, no silent stale overwrite, complete audit for accepted work, all accepted records within the deletion requirement, and sync/conflict/support metrics within preapproved thresholds. Expansion is explicitly approved at each cohort gate.
Test plan
Local persistence: Offline edits followed by process kill, device restart, low storage, and migration interruption; verify transaction boundaries, ordering, and recovery.
Sync/idempotency: Retry the same edit before/after a lost response; verify one canonical update and one audit event. Exercise intermittent network transitions, expired auth, server retry, and out-of-order responses.
Conflict: Inspector and supervisor edit concurrently; verify stale write is rejected, both candidates remain reviewable, resolution creates a new audited version, and final acceptance cannot bypass unresolved conflict.
Attachments: Capture 40 photos and a signature; test duplicate retries, partial/chunk interruption, hash mismatch, low storage, count boundary, cancellation, and finalization with missing media.
Deletion: Accept from mobile and web; exercise online, offline, powered-off, app-not-running, expired-lease, interrupted purge, duplicate receipt, and delayed server response. Verify all copies and keys are removed and the deadline is enforced by the approved platform mechanism.
Migration/compatibility: Upgrade from the current app cache, test old and new clients against additive APIs, confirm accepted cached rows are cleaned, and verify no existing server record is silently re-versioned or omitted from audit.
Security/privacy: Inspect local DB/files/backups, app logs, telemetry, and crash diagnostics for customer content; verify authorization revocation, user/device association, encrypted storage, and key destruction.
Operational: Dashboard/alert tests for queue backlog, conflict aging, missing audit writes, orphan uploads, and purge deadline risk. Do not use production customer content in test fixtures.
Decisions required before implementation
Decision
Recommendation
Why / trade-off
What starts the 24-hour deletion clock?
Define acceptance as the server's final accepted transition after the client purge receipt; gate web acceptance while an offline device lease remains active.
Gives the system an observable deletion point. It can delay supervisor acceptance when a tablet is unreachable, so verify lease/MDM behavior and business tolerance.
Audit detail and retention
Append-only server events with actor/device, server and client times, versions, event IDs, artifact references/hashes, and approved before/after details; compliance owner sets retention before M0 closes.
Supports history and deduplication; retaining field values can increase privacy and storage obligations. Do not invent a legal retention period.
Conflict handling
Strict record-level version check; no automatic merge in the pilot; preserve a conflict candidate and add a minimal supervisor compare/resolve action.
Prevents silent loss and avoids broad collaboration complexity; creates manual review work and requires a named resolver.
Offline availability
Cache only assigned active inspections; no new assignment or broad customer search offline.
Limits data exposure and keeps the offline data set bounded. Confirm assignment expiry and revocation behavior.
Photo/signature contract
Cap at 40 photos, signature separate; choose maximum bytes, allowed formats, compression, and signature representation before API freeze.
The brief gives a photo count but no payload/evidentiary limits. Resumable transfer is recommended for intermittent connectivity.
Device/security contract
Confirm tablet OS, keystore/encryption, backup exclusion, shared-device behavior, app authentication, MDM wipe, and offline lease enforcement before pilot.
A hard 24-hour physical deletion requirement cannot rest on a best-effort client task when a device may be off.
Pilot success thresholds
Set numeric thresholds for sync time/reliability, conflict volume/age, storage, support burden, and purge margin from baseline before enrollment.
Avoids arbitrary targets and makes expansion a reviewable decision.
API compatibility and audit coverage
Add new sync APIs and ensure web writes use the same version/audit guarantees before rollout; keep existing callers compatible.
Prevents a split audit history or clients bypassing version protection.
Decision log
Decision
Alternatives considered
Recommendation and rationale
Local draft authority
Keep SQLite as cache; use memory until submit; durable local workspace
Durable local workspace: it directly addresses lost notes and app/network interruption.
Concurrency
Last-write-wins; automatic field merge; strict version conflict
Strict version conflict for the pilot: safest with regulated forms and unknown field merge semantics.
Conflict resolution
Drop local copy; overwrite server; preserve isolated candidate for supervisor review
Preserve both and resolve explicitly: no silent loss or overwrite, bounded to exceptional conflicts.
Upload
Single all-or-nothing payload; independent non-resumable files; resumable artifact sessions
Resumable sessions with hashes: matches intermittent connectivity and up to 40 photos.
Acceptance cleanup
Best-effort delete after acceptance; delay final acceptance until purge evidence
Delay final acceptance until purge evidence, with a verified lease/MDM path for offline devices: gives an auditable relation between acceptance and deletion.
Rollout
Big-bang update; additive dark launch and cohort flag
Additive API first, then controlled cohorts: supports compatibility and a pause path without losing queued drafts.
gpt-6-luna/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 6.5 / 10graded blind as submission A
Covers every requested section in a readable format, but the hard problems are reasoned about weakly. The 24-hour purge has a table and a test bullet but no mechanism for a device that never learns of acceptance or has a wrong clock. Several choices risk data loss: photos upload after 'submitted' with nothing tying purge to verified upload, and the rollback flag rejects new offline submissions.
Strengths
All required sections present, with ten decisions that each carry options and a recommendation
Two-stage push with Idempotency-Key and If-Match, plus resumable photo upload
Exit criteria for each stage and per-milestone acceptance criteria
Weaknesses
Purge reasoning is thin: purge_log only, with clock skew mentioned in tests but not designed for; Q3 leaves the clock anchor open
Conflict table is muddled ('server write losing if my change was made first locally'; supervisor-supplied signature wins)
Server flag-off rejects new offline submissions, which can strand inspector work
Wi-Fi-only photo upload by default, and offline writes hard-capped at token TTL minus 30 minutes, could block inspectors in the field
RESPONSE suggests relaxing the regulatory purge rule if the assumption fails
Evidence the grader checked
§13 Rollback: 'New offline submissions rejected with a clear error'
§7 Conflict policy enum/status row: LWW by local_version ordering
§8.3 photos uploaded after the envelope is accepted; nothing gates purge on photo upload
§16 M4 server changes depend on M2 client, so the contract is built after the client
Objective checks
Files
mcode-m3/06-product-planning/RESPONSE.md
Response — Three riskiest assumptions
The plan rests on ten assumptions (see PLAN.md §1). The three below are the
riskiest because each one, if wrong, forces a redesign rather than a tuning
change. They are stated as the assumption, why it is risky, what the early
signal of it being wrong looks like, and what the plan would do about it.
1. A single inspector owns an inspection at a time (Assumption A2)
The assumption. Inspectors do not collaborate in real time on the same
inspection from two tablets. A given inspection has at most one offline writer
on a tablet at any moment; the supervisor edits from the web are a separate
concern (see risk #2).
Why it is risky. This is the foundation of the data model. The local
change log, the local_version counter, the photo ownership, the signature
ownership, and the "one envelope per inspection in flight" rule all depend on
there being a single offline writer. If two tablets can edit the same
inspection offline at the same time, the conflict policy stops being "inspector
vs supervisor" and becomes "inspector vs inspector," which is exactly the
collaborative-editor scenario the brief explicitly rejects. The audit story
also breaks: a regulator asking "who entered this note and when" needs a clean
single-writer lineage.
What would tell us it is wrong in the pilot.
Two tablets assigned to the same inspection in the assignment API.
Inspection handover in the field ("take over for me") without a server round-trip.
Multiple devices signed in by the same inspector.
Any early conflict that cannot be explained by a supervisor edit.
What the plan would do if it is wrong. Stop the pilot at Stage 1 and
re-scope. The honest fix is a CRDT for form fields, OT for ordered photo
lists, and a presence model. That is a different product, not a different
version. The right move is to either (a) add server-side single-writer
locking with a clear "claim" model, or (b) accept that handover is server-
mediated and enforce that in the assignment UI. The first is the smaller
change; the second is a process change.
2. Supervisor web edits while an inspector is offline are the exception, not the norm (Assumption A3)
The assumption. Supervisors typically open an inspection before an
inspector starts work, or after the inspector submits. Mid-visit supervisor
edits are an edge case. The plan sizes conflict UX and the conflict rate
budget against this (target < 5% of submissions).
Why it is risky. Every supervisor mid-visit edit becomes a 409 conflict
the inspector must resolve by hand. The plan intentionally rejects automatic
merging, so the inspector sees a side-by-side screen and must choose. If
supervisors routinely edit in the gap (e.g., a quality reviewer who flags
issues while the inspector is still in the basement), conflicts become a
daily tax on inspectors, the resolution UI gets exercised far more than the
pilot exit criteria anticipate, and the "no full collaborative editor" stance
starts to feel like a product gap rather than a deliberate choice. It also
stresses the audit log with conflict_resolution rows, which is fine, but
it is the kind of volume that hints the product is asking for a collab model
in disguise.
What would tell us it is wrong in the pilot.
Conflict rate above the 5% gate at any weekly checkpoint.
A pattern of conflicts on the same small set of fields (signals systematic
supervisor review).
Supervisor feedback that the conflict UX is "blocking their work."
Audit log dominated by conflict_resolution events on routine edits.
What the plan would do if it is wrong. Introduce a server-side
"supervisor notes" overlay: a separate field set that supervisors can edit
freely and that does not enter the conflict path with the inspector's
form fields. The inspector's notes and the supervisor's notes are merged at
render time. This keeps the no-collab-editor posture while removing the
operational pain. A second lever is a soft lock: when an inspection is open
on a tablet, the supervisor can read but is warned that writes will queue
behind the inspector. A third lever, more invasive, is to relax LWW for
non-overlapping fields. The first lever is the smallest viable change.
3. An offline visit is bounded, with a return-to-signal point at the end (Assumption A1)
The assumption. A tablet leaves signal at building entry and rejoins at
exit (van, parking, lobby). The visit has a known maximum duration. This is
what makes deferred photo upload viable, what makes the auth token TTL
work, and what makes the 24h purge clock line up with "the inspector has
moved on."
Why it is risky. Basements vary. Some have no signal at the surface
either, and inspectors only sync at the office or at a known Wi-Fi point
hours later. If the offline window is unbounded:
Auth tokens can expire mid-visit and block the submit (D4 hard cap is
sized for "a few hours," not "all day").
Deferred photo upload piles up across multiple visits, blowing the device
storage budget (D8) and creating a 24h-purge clock that has already
elapsed for older inspections.
The "submitted but photos not yet uploaded" window stretches, and a
supervisor opening the record sees an empty photo set with no signal that
data is in transit, which is operationally confusing.
The audit trail of an inspection can sit on the device for far longer than
anyone expected, which may itself fail the regulation that motivates the
24h rule, depending on how it is read.
What would tell us it is wrong in the pilot.
Median time from "first edit offline" to "first successful sync" above a
few hours.
Inspectors reporting that they only sync at the office or at home.
Photo queue depth growing across days on a single device.
Token-refresh failures or re-auth prompts on submit.
What the plan would do if it is wrong. Two adjustments. First, raise
the auth-token offline cap and shift to a re-auth-on-next-online rather
than a hard mid-visit cap. Second, relax the 24h-purge rule to "24h after
the device next comes online and the inspection is accepted," or, if
compliance reads the rule strictly, segment the audit log so non-customer
metadata can stay while customer data is scrubbed. Neither is free. The
cleanest path is to confirm the visit-duration assumption with the pilot
site during Stage 0 and adjust the auth and storage budgets before turning
the pilot on for real.
Why these three and not the others
The other assumptions in §1 are important but their failure modes are
narrower and have well-understood fixes: storage encryption is a swap
(D1), photo timing is a flag (D3), MDM posture is an operations problem
(§9.6), and so on. The three above change the shape of the system if they
are wrong. They are also the ones the brief is silentest on, which is why
they are the most important to validate with the product owner before M0
closes and with the pilot cohort before Stage 1 turns on.
mcode-m3/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
mcode-m3/06-product-planning/PLAN.md
Plan — Offline inspection sync for tablet app
Source product input: BRIEF.md (only). This plan makes assumptions explicit
and lists the decisions that must be settled before implementation.
1. Assumptions
These are the working assumptions used to shape the plan. Each is restated as a
decision in §14 with a recommendation.
#
Assumption
Source / basis
A1
An "offline inspection" is bounded to the duration of one site visit: a tablet leaves signal on entry and rejoins on exit (van, parking, lobby).
"basements with no signal" + "connectivity may return intermittently"
A2
Inspectors do not collaborate in real time on the same inspection from two tablets.
Not stated; absence of multi-inspector requirement and product's "no full collaborative editor" stance.
A3
Supervisor web edits are infrequent and typically happen before an inspection opens on the tablet or after it is submitted. Mid-visit supervisor edits are the edge case.
"Supervisors can edit the same inspection from the web" without more.
A4
Average per-inspection payload: ~30 form fields, ≤40 photos at ~3–6 MB each, 1 signature image, free-text notes. Worst-case ≈ 300 MB local.
Stated photo cap; common tablet camera defaults.
A5
The tablet fleet is company-managed (MDM/DEP), single primary user per device, OS-level encryption at rest is on, and the app is forced-managed.
"company-managed tablets"; standard posture.
A6
"Customer data" in the 24h rule includes the inspection record, notes, photos, signature, and derived metadata; it excludes server-side IDs, sync log, and audit copies.
Reasonable reading of the regulation phrasing.
A7
Existing REST APIs already return integer recordVersion on inspection reads and accept it on writes (If-Match style).
"Existing server APIs use REST with integer record versions."
A8
The pilot is a single region / customer / team, gated by a server-side feature flag and an MDM config, not a separate code branch.
"Product wants a pilot before broad rollout."
A9
Photo upload can be deferred relative to form + signature: form must reach the server to mark an inspection "submitted"; photos can drain afterwards.
Common offline pattern; reduces required online window.
A10
Tablets run a recent OS version that allows background fetch and constrained network access.
"company-managed tablets"; standard fleet.
2. Goals and non-goals
2.1 Goals
G1. Inspectors can open an assigned inspection, fill forms, take up to 40 photos, and capture a signature with no network.
G2. Submitted inspections are durable across app crash, OS kill, and forced reboot.
G3. Sync is automatic and opportunistic; an inspector does not have to "press sync."
G4. Server remains authoritative; conflicts are detected, surfaced, and resolved by a defined policy.
G5. Local customer data is removed within 24 h of an inspection being marked accepted on the server.
G6. A regulator-grade audit history exists for every record field, photo, and signature.
G7. Pilot can be enabled for a subset of users and rolled back without a client release.
2.2 Non-goals
N1. Real-time multi-user collaborative editing (no operational transform, no CRDT, no live cursors).
N2. Inspector-to-inspector chat or handoff of an in-progress inspection.
N3. Offline supervisor workflows. Supervisors continue to require a connection; offline shows a banner.
N4. General-purpose offline mode for unrelated modules of the app. Scope is inspection write path only.
N5. A new backend service. We extend existing REST endpoints and a single new endpoint family.
N6. Server-side conflict resolution UI for inspectors. Conflicts surface on the tablet, not in a web tool.
3. User flow
3.1 Inspector flow (offline-first)
Inspector opens the app while online. App pulls the assigned inspection list (delta since last sync) and pins it locally.
Inspector taps an inspection. App verifies the record and attachments are fully local; otherwise shows a "needs sync" indicator and queues a fetch.
Form opens from local SQLite. Status banner: "Offline — saving locally."
Inspector edits fields. Each edit writes a row in local_change (append-only log) tied to the current local version.
Photos are captured to the app's sandboxed storage; metadata is added to local_change with a content-hash and byte size; the binary is uploaded separately.
Signature is captured as an SVG (or PNG fallback) and added to local_change.
Inspector taps Submit. App:
Marks the local record pending_submit (still in sync queue even if online).
Promotes changes to a "submission" envelope: form patch + ordered list of photo IDs + signature ID.
On any next network event, pushes the submission.
After successful form push, photos drain in the background. Status changes syncing_photos → synced.
Once synced reaches the server, the record transitions to submitted server-side and surfaces in supervisor view.
3.2 Inspector flow (intermittent connectivity)
The app listens for network changes. On every transition to "connected," it:
Pulls deltas for assigned inspections and any open record.
Drains pending submissions oldest-first.
Drains pending photo uploads, newest-photo-first skipped if size budget exceeded.
Resumes any failed transfer from a checkpoint.
3.3 Supervisor flow (web, mostly online)
Continues to work against the server.
If supervisor opens an inspection that an inspector is currently editing offline, the supervisor sees a "device last seen offline; edits will reconcile on sync" banner with the device's last-known local version.
Supervisor saves with their own recordVersion. The inspector's later push hits a version conflict; the inspector resolves per §6.
3.4 Conflict UX (inspector)
On opening a record post-sync, if the server version advanced while the inspector was offline, the app shows a side-by-side conflict view:
The form snapshot in inspection.body_json is the source of truth for rendering. local_change is the audit / replay source.
local_version is monotonic per inspection. Every read of body_json is paired with the current local_version so editors can detect clobber.
Photo binaries live on the filesystem; the row tracks the path and content hash. The hash is also sent to the server for de-duplication.
5.3 Storage budget
Per-device cap on pending inspection size (configurable, default 1 GB). The app refuses to start a new offline inspection when the cap would be exceeded and offers "submit or purge" actions.
6. Sync protocol
6.1 Pull (server → client)
GET /inspections/assigned?since=<cursor>&limit=N — list delta.
GET /inspections/{id} — full record with recordVersion, ETag, and links to expected attachments.
GET /inspections/{id}/attachments?ids=... — resumable attachment download (range requests) for any blob the client does not already have by content hash.
All GETs are idempotent. Cursor is server-issued; client persists last cursor per device.
Body: form patch (JSON Patch or simple {field: value} diff), ordered list of photo IDs (client UUIDs, content hashes, sizes), signature ID, local audit tail hash.
Server response:
200 with new recordVersion → mark envelope accepted.
409 with { serverVersion, conflicts: [{field, mine, theirs}] } → mark envelope conflicted; client renders the conflict UI.
412 (precondition failed) → fetch latest, recategorize as conflict, recurse.
429 with Retry-After → schedule retry.
Stage B — photo uploads (deferred)
POST /inspections/{id}/photos with multipart, Idempotency-Key, Content-SHA256.
Resumable: if a transfer is interrupted, the next call resumes from bytes_sent.
Server may return 200 immediately if the hash matches an existing photo on the inspection (dedupe).
6.3 Push ordering
Stage A for an inspection must complete before Stage B for the same inspection.
Across inspections: oldest first by envelope.created_at.
Photos: largest first within an inspection (frees local space faster) — but the server may override.
6.4 Idempotency
Every write carries Idempotency-Key and is stored on the server with TTL ≥ 7 days. A retried envelope with the same key returns the original outcome.
The client's local_change.client_change_id is the idempotency key.
6.5 Bandwidth and battery controls
Sync runs only on Wi-Fi by default for photo uploads unless the user toggles "use cellular" for a session.
Sync is throttled (configurable concurrency, default 2 photo uploads in flight).
Sync backs off on low battery (<20%) for non-critical transfers.
6.6 Server assumptions
Server stores recordVersion per inspection and increments on every accepted write.
Server treats Idempotency-Key as a write-once key.
Server returns 409 conflicts with structured payloads; it does not silently overwrite.
7. Conflict policy
The product explicitly rejects a full collaborative editor. We use detection + explicit per-field resolution on the device, with the server as authority.
Field type
Policy
Notes
Free-text form fields
Per-field: keep mine / keep theirs
No automatic merge.
Enum / status fields
Last-writer-wins, with the server write losing if my change was made first locally (LWW by local_version ordering)
Simpler; conflicts are rare.
Photos
Additive; both kept
List semantics, not field semantics.
Signature
Server wins if a supervisor uploaded one; otherwise keep mine
Regulation: the captured signature is the customer's.
Audit metadata (timestamps, device id)
Server always wins
Don't fight on system data.
State transitions (open → submitted → accepted)
Server is the only writer for accepted
Inspector never marks accepted.
The conflict UI (§3.4) is mandatory: the inspector must resolve before the resubmission is queued. There is no automatic re-submit after a conflict. Every resolution choice is written to local_change with kind=conflict_resolution for the audit trail.
8. Attachments
8.1 Photos
Capture: stream directly to the app's encrypted sandboxed directory; never to the OS gallery.
Naming: inspection_<id>/photo_<uuid>.jpg plus a sidecar .json with EXIF minus GPS, capture time, hash.
Strip GPS at capture time if device policy allows; otherwise flag for the post-accept purge.
Compression: client-side downscale to a configured max edge (default 2048 px) and quality (default 0.8) before storing. Originals are not kept.
Thumbnail generated at capture (256 px, low quality) for fast list rendering.
8.2 Signature
SVG first; PNG fallback for older firmware.
Stored as a separate record, not embedded in form JSON, so it can be replaced or audited independently.
8.3 Uploads
Photos are uploaded after the form submission envelope accepts. This decouples the regulatory "submitted" state from photo transfer success and lets a flaky link still deliver the data.
A photo is uploaded only when the server returns success. Until then, the local copy is the source of truth.
Storage pressure: if local storage drops below a threshold, the app uploads the largest pending photos first.
8.4 Other attachments
Out of scope. No PDFs, videos, or voice memos in v1. If added later, they use the same photo pipeline with a different MIME profile.
9. Security
9.1 At rest
Rely on OS-level full-disk encryption (assumption A5).
App sandbox on iOS / Android is non-negotiable; no world-readable files.
Optional: an additional per-app key wraps the SQLite file via SQLCipher. Decision D1 (see §14). Default off unless compliance requires it.
9.2 In transit
TLS 1.2+ only; certificate pinning for the inspection API host. Pin rotation requires a server-side pin-overlap window.
Reject any non-TLS connection, including captive portals that downgrade.
9.3 Authentication
Same identity as the existing tablet app. Token is refreshed online; offline use is bounded by the existing token TTL.
Offline window must not exceed token TTL minus a safety margin; if it does, the app requires re-auth before allowing edits and surfaces a clear message.
9.4 Authorization
Server-side: the same checks as today. A new "write offline envelope" permission is added; existing users in the pilot group receive it; others don't.
9.5 Data minimization
Strip EXIF GPS unless policy says otherwise.
Do not log raw field values to client logs; only record IDs and counts.
The audit log on the device mirrors the server audit log and is itself scrubbed on the same 24h clock.
9.6 MDM posture
App is managed; remote wipe of the app's data is supported via the standard managed-app remove signal.
"Block screenshot of inspection screen" is enabled on iOS where possible, given signatures and customer data appear on screen.
10. Migration and backfill
10.1 Schema migration
Migration scripts are versioned; the new tables are additive. Existing rows in the read-through cache are forward-compatible.
A migration test seeds a v0 database (current schema) and asserts v1 upgrades cleanly with no data loss and the new tables empty.
10.2 Backfill of existing inspections
No data migration is required for inspections already on the server. They stay on the web/online path.
An inspection opened for the first time on a tablet after rollout is downloaded with the new schema, including body_json populated from the existing API response.
10.3 Pilot user backfill
Pilot users get the new build via MDM. Their local cache is rebuilt on first launch (full pull, not delta) so the new tables and state machine start clean.
A "reset local cache" admin action is available for support.
10.4 Deprecation
The old read-through cache path is kept for read-only flows outside the inspection write path. Removal is a later project.
11. Observability
11.1 Client-side
Every sync run writes a sync_run row.
Emit (locally, with sampling) metrics: time-to-first-paint for an inspection, time between edits, time from "online" to "submitted accepted," photo queue depth, retry counts, conflict counts.
Client logs are scrubbed; no PII in logs.
11.2 Server-side
New counters and traces: offline_envelope_received_total, offline_envelope_conflict_total, offline_photo_upload_bytes_total, offline_idempotent_replay_total.
Trace each envelope end-to-end via Idempotency-Key.
Dashboards split by pilot cohort.
11.3 Alerts
Pilot-only alerts: spike in 409 conflicts, spike in envelope failure rate, drop in submission throughput, surge in client crash reports on inspection screen.
An alert on accepted_but_photos_not_uploaded is a first-class signal; both must converge.
11.4 Audit
The audit log captures who, what, when, from which device, on which recordVersion. It is replicated to long-term storage and is regulator-readable.
12. Staged rollout
Each stage has a go / no-go gate. Rollback path is defined per stage.
12.1 Stage 0 — internal dogfood
Audience: engineers and a small QA team, synthetic data.
Goal: prove the state machine, conflict UI, and idempotency on the wire.
Exit criteria: 200 successful synthetic submissions offline→online; 50 synthetic conflicts resolved; no data loss across app kill mid-submit.
12.2 Stage 1 — closed pilot
Audience: one customer or one region, ≤ 25 inspectors, with their supervisors still on the web tool.
Goal: validate the inspector experience in real basements.
Flags: server-side OFFLINE_INSPECTION_PILOT and an MDM configuration profile. Off by default.
Exit criteria: ≥ 90% pilot inspections complete without manual supervisor help; no unresolved data loss; conflict rate < 5% of submissions; 24h purge verified on at least 20 accepted inspections.
12.3 Stage 2 — expanded pilot
Audience: 1–3 more customers, up to 200 inspectors.
Goal: stress on bandwidth variability and concurrent supervisor edits.
Exit criteria: same as Stage 1 plus supervisor-edit-during-offline scenario tested with at least 10 real occurrences and no customer-visible data corruption.
12.4 Stage 3 — general availability
Server flag flipped to on for all customers. MDM config no longer required to enable.
Rollback is the same flag flip.
13. Rollback
Layer
Rollback action
Data impact
Server flag
Flip OFFLINE_INSPECTION_PILOT off
New offline submissions rejected with a clear error. In-flight envelopes are allowed to drain.
App build
MDM removal of the offline-enabled build
Tablets re-downgrade to the previous build; new edits become online-only.
Local data
App detects flag-off and stops opening offline records; existing cached records remain read-only.
No data loss; offline records become ghost rows.
Server endpoint
Disable new endpoints behind feature flag
Clients see 404 on submit; UX is "this inspection can only be edited online."
Worst case
App + server flag both off, all clients downgraded
The system is back to pre-rollout behavior with a clean cutoff.
No destructive server action is required to roll back. Because the server is authoritative and the offline envelope uses Idempotency-Key, no partially-applied state is created during rollback.
14. Decisions to make before implementation
Each decision has a recommendation. The first three are the highest-priority and feed directly into the conflict policy and the data model.
Recommendation: (a) OS-only for v1. Tablets are managed and FDE is on. Add (b) if a security review or regulation demands it; the schema is ready. (c) is overkill and complicates resume.
D2. Conflict resolution model
Options: (A) Per-field explicit resolution only, (B) Automatic LWW for non-overlapping fields + explicit for overlaps, (C) CRDT-style merge for text fields, (D) Server-side merge tool.
Recommendation: (A). Matches the "no full collaborative editor" stance. Keep the audit trail simple. Revisit if pilot shows >10% of submissions hit conflicts.
D3. Photo upload timing
Options: (1) Photos ship in the same envelope, (2) Photos upload only after form submission is accepted, (3) Photos upload as soon as the device is online, regardless of form.
Recommendation: (2). Decouples regulatory "submitted" from photo transfer. Lets the inspector leave the basement as soon as the form is uploaded. Worst case: accepted form with photos not yet uploaded is visible in the supervisor UI as a banner.
D4. Token TTL and offline window
Options: extend token TTL to cover a worst-case visit, shorten offline capability, or require re-auth after N hours.
Recommendation: keep current TTL, hard-cap offline writes at TTL - 30 min with a visible timer, and require re-auth to extend. Aligns with security without server changes.
D5. Pilot gating mechanism
Options: server-side feature flag, MDM config, or both.
Recommendation: both. Server flag for the contract, MDM config for the app's UI, kept in sync via a capabilities endpoint the app calls on launch.
D6. Cellular data policy
Options: Wi-Fi only for photos, Wi-Fi only for everything, allow cellular with opt-in.
Recommendation: Wi-Fi-only for photos by default, cellular opt-in per session, form + signature can use any network. Document the cost implications for the customer.
D7. EXIF / GPS handling
Options: keep, strip GPS only, strip all EXIF, keep EXIF but redact GPS server-side.
Recommendation: strip all EXIF at capture. Storage and privacy win; audit timestamps come from the app, not the file.
D8. Storage budget per device
Options: hard cap at 1 GB, configurable per cohort, no cap.
Recommendation: configurable per cohort, default 1 GB. Inspectors in the pilot are told the cap; the app enforces it.
D9. Server dedup of photos
Options: dedup by content hash globally, per inspection only, no dedup.
Recommendation: per inspection, by content hash. Cheap server-side, removes accidental duplicates, no cross-customer storage cost.
D10. Audit retention
Options: retain device audit log only until purge, retain server audit for N years, replicate to WORM storage.
Recommendation: server audit is retained per the existing records-retention policy (decision belongs to compliance — flag it). Client audit is purged on the same 24h clock as data.
Idempotency: replaying the same envelope key returns the same server result; double-tap on submit does not double-write.
Version conflict: a synthetic supervisor edit between envelope creation and submit triggers the conflict UI and audit row.
Resume: kill the app mid-photo-upload; on relaunch, the photo resumes from bytes_sent.
15.3 Network conditions
A network lab profile (network link conditioner or app-level proxy) covering:
No signal for the full visit, then Wi-Fi only.
Flapping connection every 30 s.
Captive portal that returns HTTP 200 to a body that is not JSON.
TLS interception (must be rejected by pinning).
High-latency, low-bandwidth (GPRS-equivalent).
15.4 Device conditions
Storage pressure (≥95% full).
Low battery.
OS upgrade mid-visit.
Forced reboot.
Two backgrounded sync runs racing.
15.5 Compliance
Audit log integrity check (tamper-evident hash chain over local_change rows).
24h purge verification on a fleet of test devices with a wall-clock advance.
Privacy review of every field, photo, and signature stored at any lifecycle stage.
15.6 Performance budget
Cold start to first inspection: target under the existing app's current number plus a documented overhead.
Edit-to-disk: target < 50 ms on mid-tier tablet hardware.
Photo capture to local write: target < 200 ms.
Sync drain of a 40-photo, fully-pending inspection on good Wi-Fi: target bounded by bandwidth; tracked, not asserted.
15.7 Security
Pen test of the new endpoints, including the Idempotency-Key semantics.
Token-expiry mid-visit test.
Screenshot-block enforcement on iOS.
16. Milestone sequence
Milestones are ordered; later milestones can start once their dependencies are met. Effort is described as scope, not headcount, to avoid unjustified staffing estimates.
M0. Decisions and design freeze
Dependencies: none.
Work: ratify D1–D10, finalize data model, finalize conflict UX mocks, finalize server endpoint contract.
Acceptance: this document is approved, mockups are clickable, the server team has an API spec merged.
M1. Local data + state machine
Dependencies: M0.
Work: SQLite migrations, inspection state column, local_change, envelope, photo, signature tables, state machine with unit tests.
Acceptance: 100% state transition coverage, migration test from v0 → v1 passes, replay test passes.
M2. Form-only offline (no photos, no signature)
Dependencies: M1.
Work: form editing offline, push on connect, version conflict path on the form fields, audit log.
Acceptance: synthetic end-to-end test passes; conflict UI renders and resolves correctly.
Acceptance: 40-photo inspection uploads and resumes correctly under flaky network; signature upload is atomic.
M4. Server-side changes
Dependencies: M0 (API spec), M2 (form).
Work: recordVersion semantics, If-Match and Idempotency-Key middleware, conflict response shape, photo dedup, new endpoints.
Acceptance: contract tests pass against the M2 client; load test of 10x pilot volume shows acceptable p95.
M5. Observability and audit
Dependencies: M3, M4.
Work: client metrics, server metrics, audit hash chain, dashboards, alerts.
Acceptance: every metric in §11.2 has a dashboard and a documented alert; audit chain verification tool exists.
M6. Pilot enablement
Dependencies: M3, M4, M5.
Work: feature flag, MDM config, capabilities endpoint, pilot runbook, support tooling (reset local cache, force resync, view device state).
Acceptance: dogfood (Stage 0) exits its go/no-go gate.
M7. Closed pilot (Stage 1)
Dependencies: M6.
Work: pilot cohort onboarding, training, daily ops check-in.
Acceptance: Stage 1 exit criteria met.
M8. Expanded pilot (Stage 2)
Dependencies: M7.
Work: more cohorts, supervisor-edit scenarios.
Acceptance: Stage 2 exit criteria met.
M9. GA (Stage 3)
Dependencies: M8.
Work: flip flag, communication, deprecate any temporary workarounds.
Acceptance: flag-on cohort = all customers; rollback drill run successfully.
M10. Post-GA hardening
Dependencies: M9.
Work: revisit conflict policy if real rate > threshold, revisit encryption decision, revisit photo size policy, document known limits.
Acceptance: known-limits doc is published; backlog is groomed.
17. Open questions for the product owner
These are not blockers for M0, but they shape the UX.
Q1. What is the maximum acceptable time from "form submitted" to "all photos visible to the supervisor"? This drives photo upload prioritization.
Q2. Can the supervisor web tool add a "currently being edited offline" banner sourced from a heartbeat, or is that out of scope for v1?
Q3. Is the 24h purge measured from server-acceptance timestamp or from device-side clock? Server timestamp is the only honest answer; confirm.
Q4. Are there any inspections that must be 100% online (e.g., for legal reasons in a given jurisdiction)? If yes, the app must be able to block offline open for those records.
Q5. What is the contractual SLA for "submitted but photos not yet uploaded" visibility, if any?
18. Summary
The plan is an offline-first write path on the tablet, with explicit conflict
resolution, deferred photo upload, server-authoritative versioning via the
existing integer recordVersion, and a strict 24h local-data purge. It is
delivered in additive migrations, gated by a feature flag, and rolled out
through three pilot stages before general availability. The next move is to
ratify the ten decisions in §14 and freeze the API contract.
mcode-m3/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.
Evaluation 5.75 / 10graded blind as submission I
Short and covers the headings, but the hard problems are handled shallowly. Purge waits until the device learns of acceptance, yet the tests claim to verify the 24-hour rule for an offline device. There is no encryption of local data at rest. Supervisor-wins drops the inspector's on-site value to an audit row and a toast. Pilot thresholds and cohort sizes are invented without justification.
Strengths
Concise, with idempotent changes, a purge-ack endpoint and server-authoritative timestamps
Kill switch keeps the outbox draining
Ten decisions, each with a recommendation
Weaknesses
No mechanism for a device offline at or after acceptance; purge happens 'on next launch'
Security covers token storage and app-private files but not encryption of SQLite or photos
No separate server milestone; acceptance criteria are thin
Evidence the grader checked
§3 step 5–6: purge-after delivered on next device sync
§9 Security: no at-rest data encryption
§7 'local field loses (server wins)... non-blocking toast'
§12 S1 gate metrics
Objective checks
Files
mimo/06-product-planning/RESPONSE.md
RESPONSE — three riskiest assumptions
Offline auth stays usable for a full shift (A1/D5). If tokens die or the server demands fresh auth while the inspector is in a basement, every other offline feature is unreachable. This is the single point of failure before capture even starts. Mitigation: device-bound cached credentials gated by a daily biometric/passcode unlock (D5); validate on real devices in S0 before building anything else.
Conflict policy of "server (supervisor) wins on same-field edit" is acceptable to inspectors and compliance (A8/D4). If field staff cannot accept losing a locally edited value (even with an audit copy and a toast), the product needs merge UI or offline edit leases — a large scope increase the brief explicitly forbids. Mitigation: put the conflict dialog in front of pilot users in S0 and treat any "I lost my notes" complaint as a gate failure, not a support ticket.
The 24h post-acceptance local purge requirement can be met with a purge receipt (hash list) rather than a local audit replica (A4/D6). If regulators require on-device audit history after acceptance, purge and retention collide, the data model changes, and mobile storage may never be demonstrably clean. Mitigation: confirm with compliance before M4; the receipt design is a cheap extension point but the underlying deletion is not reversible, so this must be locked before the first real acceptance in the pilot.
mimo/06-product-planning/BRIEF.md
Offline inspection brief
Inspectors use company-managed tablets in basements with no signal. Today, a
lost connection after opening an inspection can discard notes. They need to open
assigned inspections, fill forms, take up to 40 photos, and collect a signature
offline. Connectivity may return intermittently. Supervisors can edit the same
inspection from the web. Regulations require an audit history and removal of
local customer data within 24 hours after an inspection is accepted. Existing
server APIs use REST with integer record versions. The mobile app has SQLite but
currently treats it as a read-through cache. Product wants a pilot before broad
rollout and does not want a full collaborative editor.
mimo/06-product-planning/PLAN.md
Plan: Offline work for tablet building-inspection app
1. Assumptions (explicit)
#
Assumption
If wrong
A1
Company-managed tablets have a local auth session (token or cert) that remains valid for a full offline shift (≥ 24h)
Need offline-issued short-lived credentials or a device unlock flow
A2
Assigned inspections can be listed and fully downloaded before signal is lost; inspectors know which assignment they will work
Need an always-cached "my queue" and forced prefetch on assignment
A3
Acceptance (supervisor sign-off) is a server-side state transition; the tablet only submits, it does not accept
Acceptance path also needs offline rules; scope grows
A4
"Removal of local customer data within 24h after acceptance" means: after sync confirms accepted, the device may purge that inspection; no need to keep a local audit replica
Must retain a minimal audit stub locally and prove retention to auditors
A5
Forms are structured (schema-versioned fields + free-text notes), not freeform documents
Richer offline editing model required
A6
Photo metadata (time, inspector, inspection id) is the compliance artifact; exact original bytes may be recompressed for transfer
Must keep original bytes bit-for-bit
A7
Integer record versions on the server are monotonically increasing per inspection record and can be used as sync tokens
Need a richer revision/history API
A8
One inspector owns field edits on a given inspection during a session; supervisor edits are rare and can lose to "last writer with version check" rather than merge
True multi-author field-level merge is required (out of scope per product)
A9
Pilot devices can receive app updates and server feature flags; pilot users tolerate occasional re-sync prompts
No remote control channel; bake everything into releases
2. Goals / non-goals
Goals
Open assigned inspections, fill forms, attach up to 40 photos, capture a signature with zero connectivity.
Survive intermittent connectivity with automatic retry; never silently discard user-entered data.
Keep a tamper-evident audit history of field-level changes and sync events.
Purge customer data from the device within 24h of server-side acceptance.
Offline editing by supervisors on the web (web stays online-only).
Cross-device handoff of an in-progress offline inspection.
3. User flow
Prep (online): Inspector signs in, opens My assignments. App downloads each assigned inspection (schema + current answers + photo list + version) and marks it local-ready. Badge: "Available offline".
Site (offline): Inspector opens a local-ready inspection. Forms save to SQLite on every field commit. Photos write to app-private storage with a local blob id; thumbnails cached. Signature captured as vector + raster; hash recorded.
Intermittent sync: When the network returns (app foreground, or background fetch every N minutes when permitted), the app drains a per-inspection outbox: metadata/answers first, then attachments in priority order (signature > small fields > photos). Progress UI shows pending count. Retry with exponential backoff + jitter; never drop an outbox entry without either success or explicit user confirmation of abandon.
Submit offline: Inspector marks the inspection ready-to-submit. If offline, it queues as the final outbox item. On successful submit, server returns new version + submitted state.
Acceptance (supervisor, web): Supervisor reviews, may edit (online). Server bumps version and records audit. When state becomes accepted, the next device sync response includes accepted + purge-after timestamp (accepted_at + 24h).
Purge: App deletes local answers, photos, signature, and derived thumbnails for that inspection at purge time (or on next launch if device was off), keeps only a purge receipt (inspection id, purge time, content hashes) unless A4 says otherwise, and reports purge to the server for compliance.
4. Scope boundaries
Offline is available only for inspections that were successfully opened/downloaded while online at least once ("local-ready set").
Edits are whole-inspection last-writer-wins at field granularity with version precondition; no merge UI in v1 (conflict surface is a dialog, see §7).
Photos: up to 40 per inspection, compressed for upload (configurable max edge / JPEG quality), originals retained locally until purge.
Signature: one inspector signature offline; countersign (if any) is online-only in v1.
No offline export/email of reports; report PDF generation stays server-side.
Blobs in files/attachments/<sha256> (content-addressed). DB stores only hashes + paths. All writes go through a single repository layer so purge and audit cannot be bypassed.
local_rev is a per-device monotonic counter for idempotent outbox identity; the server still owns server_version.
6. Sync protocol
Keep REST; extend, don't replace.
PullGET /inspections/{id}?since_version=v → record + version + state. Also GET /assignments for queue refresh (online only).
PushPOST /inspections/{id}/changes with:
base_version (precondition),
client_change_id (UUID, idempotency key),
field upserts (only dirty fields) + client_rev,
optional op=submit.
Response: 200 new server_version, or 409 with server field snapshot for conflict handling, or 422 schema/validation error (do not retry blindly — surface to user).
AttachmentsPOST /inspections/{id}/attachments (multipart or resumable upload session for photos; chunk size ~1–2 MB). Server returns server_id. Upload is content-hash idempotent. Signature uploads first and is required before submit succeeds.
Purge ackPOST /inspections/{id}/purge-ack with device id, purge time, list of deleted hashes.
Ordering per inspection: answer upserts → attachment uploads → submit → (later) purge-ack. Across inspections, outbox is parallelizable but bounded (2 workers) to protect radio/battery.
Idempotency: every push carries client_change_id; server stores it and returns the original result on replay (covers timeout-after-commit).
Clock skew: server timestamps are authoritative for accepted_at / purge_after; device uses them verbatim.
7. Conflict policy
Product decision: no collaborative editor. Policy:
Every push requires base_version == server_version, else 409.
On 409, app compares dirty local fields vs server snapshot:
If disjoint field sets → auto-rebase: retry push with new base_version (still recorded in audit).
If same field edited → local field loses (server wins) for supervisor-originated changes; show a non-blocking toast listing overwritten fields, and keep the local discarded value in audit_events (not as an editable answer).
If local has submit pending and server moved past submitted/accepted → drop local submit, adopt server state.
Photo/signature blobs never conflict (append-only by hash); duplicates deduped server-side.
Rationale: supervisors are the reviewers; their online corrections are authoritative. Field-level LWW avoids silent data loss better than whole-record LWW.
8. Attachments
Cap 40 photos/inspection (enforced locally and server-side). Show remaining count.
Local pipeline: capture → downscale preview → store original → enqueue upload of transfer encoding (config: max edge 2048, JPEG q≈80) plus EXIF time. Strip precise GPS by default (policy flag).
Resumable uploads; cancel/retry UI; upload only on Wi-Fi toggle (default on for company-managed devices).
Signature stored as PNG/SVG + device-local cert binding (see §9). Hash of signature payload written to audit chain before first upload attempt.
9. Security
Auth tokens in platform keystore/Keychain; biometric/passcode device unlock required before first offline edit of the day (gate for offline use of cached credentials).
App-private storage; no customer data on shared storage. Screen capture discouraged via OS flags where available (not compliance-critical).
All sync traffic TLS 1.2+; certificate pinning optional for pilot (decide via D1).
Field-level values never logged; logs carry ids/hashes only.
Remote wipe / selective purge command path (MDM or server "revoke inspection" push) — MDM wipe covers lost devices; in-app purge covers retention.
Signature integrity: hash of signature image + inspection id + schema version chained into audit; server verifies hash on submit.
Device: migrate answers table from cache semantics to durable outbox-backed store (local_rev, sync_state columns; backfill local_rev=0). Migration is transactional; on failure app falls back to read-only cache mode and forces resync rather than risking corruption.
Existing cached inspections are re-downloaded once (treat as not local-ready until re-fetched) — simplest correct backfill.
Audit chain starts at migration time; a server-side audit_created root event anchors the chain per inspection.
11. Observability
Client metrics (aggregated, no PII): outbox depth/age, sync success/409/422 rates, attachment upload retries/time-to-drain, purge latency after accepted_at, crash-free sessions on pilot builds.
Server metrics: 409 rate by field, idempotent replays, purge-ack lag, attachment integrity mismatches.
Per-inspection audit chain export endpoint for regulators (server-side; device contributes events via sync).
Alerting: purge-ack lag > 20h in pilot → page on-call (compliance clock).
2 full offline-to-accept-to-purge cycles per device; zero silent data loss
S1
Pilot: 5–15 inspectors, one region, real jobs
≥ 95% sessions complete offline path without support ticket; purge-ack lag p95 < 12h; 409 auto-rebase success ≥ 90% of conflicts
S2
Expand regions + supervisor conflict tooling polish
S1 metrics hold for 2 weeks of real volume
S3
Broad rollout
No open Sev-1/2 from S2
Feature flag offline_mode (server-driven) plus minimum app version flag. Kill switch flips devices to read-only cache behavior on next connect.
13. Rollback
Flag off: devices stop accepting offline edits; pending outboxes still drain (never discard). New edits require connectivity.
App release rollback: schema migrations written backward-compatible within one minor (additive columns only for 2 releases) so store rollback doesn't brick local DB.
Server rollback: v2 routes are additive; v1 clients unaffected. Idempotency store retained.
Data rollback: purge is destructive by design — not reversible; rollback of the feature does not resurrect purged device data (server remains system of record).
Declare explicitly: accept loss of in-flight offline edits only via user-mediated "abandon changes" — no automatic rollback path drops outbox.
AC: S1 gate metrics met or explicit product sign-off to iterate.
Dependencies: M1→M2→M3→M4 is the critical path; M5 can overlap late M4 (flags/observability); M6 is strictly gated.
16. Decisions required before implementation (with recommendations)
ID
Decision
Recommendation
Why
D1
TLS pinning on/off
Off for pilot, measure; decide before S2
Pinning complicates enterprise proxy debugging; pilot risk is low on managed networks
D2
Photo transfer encoding vs originals
Upload compressed derivative + hashes of original; keep original until purge
Bandwidth in basements/street 4G; regulators accept hash-anchored originals
D3
GPS in EXIF
Strip by default; site id is the location record
Avoid precise home location retention; simpler purge surface
D4
Conflict: who wins on same-field edit
Server (supervisor) wins; local discarded value kept in audit only
Matches "not a collaborative editor"; reviewers are authoritative
D5
Offline auth lifetime
Require daily unlock (biometric/passcode) to use cached creds; session ≤ 7 days device-bound
Limits lost-device exposure while keeping full-shift offline
D6
Purge receipt contents
Store id, times, hash list only — no field values
Satisfies A4 with minimal residual risk; easy to extend if auditors demand more
D7
Attachment upload radio policy
Wi-Fi preferred, cellular allowed after 30 min backlog
Pilot devices are cellular-enabled; don't strand evidence
D8
Schema evolution while offline
Pin schema_version at download; edits with old schema still accepted if field ids stable; force re-download if schema breaking: true
Avoids mid-form surprise; server remains flexible
D9
Signature acceptance offline
Capture full signature packet offline; mark submitted only after signature hash reaches server
Signature is the legal artifact; don't let submit succeed without it
D10
Kill-switch behavior
Read-only cache mode, outbox drains, no new offline edits
Never destroy user work when disabling the feature
17. Open questions (explicitly deferred)
Countersignature workflow and whether it needs offline support (assume no for pilot).
Cross-inspection batching for sites with many units.
Long-term device audit retention if A4 is later rejected by compliance.
mimo/06-product-planning/TASK.md
Task 06 — Product and engineering planning: offline inspection sync
Create PLAN.md for adding offline work to a tablet-based building inspection
app. Use BRIEF.md as the only product input.
The plan should make assumptions explicit and cover user flow, scope boundaries,
local data model, sync protocol, conflict policy, attachments, security,
migration/backfill, observability, staged rollout, rollback, and testing. Include
a milestone sequence with dependencies and acceptance criteria. Identify the
decisions that must be made before implementation and offer a recommendation for
each. Avoid fake dates and unjustified staffing estimates.
Write RESPONSE.md summarizing the three riskiest assumptions. Work only in this
directory.