QuickBooks Online Local Backup (QBO-Archive)
Section titled “QuickBooks Online Local Backup (QBO-Archive)”Purpose
Section titled “Purpose”Create a deterministic, local backup of an entire QuickBooks Online company using the Intuit Developer API.
The backup is read-only and intended solely for disaster recovery, historical preservation, migration assistance, and offline inspection.
The output is a single JSON file.
No attempt is made to restore the data automatically. A future restore tool may consume this JSON.
The backup system should be:
-
Complete
-
Deterministic
-
Repeatable
-
Idempotent
-
Easy to validate
-
Human inspectable
-
AI inspectable
-
Independent of Intuit
The system should produce exactly one file:
D:\FSS\Accounting\QuickBooks\QBO-archive.jsonThe file is overwritten each execution.
Historical versions are retained by Kopia.
Technology Stack
Section titled “Technology Stack”Python
Package manager
- uv
Libraries
-
httpx
-
pydantic
-
orjson
-
loguru
Reuse the existing OAuth implementation already used by the transaction-processing project.
No third-party QuickBooks wrappers unless they provide a clear long-term maintenance advantage.
Project Structure
Section titled “Project Structure”qbo-archive/
pyproject.toml
src/
archive.py api.py auth.py config.py models.py pagination.py serializer.py main.pyDesign Principles
Section titled “Design Principles”1. Read Everything
Section titled “1. Read Everything”The backup should retrieve every supported entity available through the QBO REST API that the authenticated company exposes.
Never back up only reports.
Always back up the underlying source objects.
2. Preserve Raw Data
Section titled “2. Preserve Raw Data”Do not normalize fields.
Do not rename fields.
Do not remove unknown properties.
The backup should preserve the API response exactly whenever practical.
If additional metadata is added by the backup system, place it outside the raw API objects.
3. Stable Ordering
Section titled “3. Stable Ordering”Objects must always be sorted before writing.
Recommended ordering:
-
entity type
-
primary identifier
-
creation date
-
update timestamp
Stable ordering makes Git diffs and backup comparisons much easier.
4. Deterministic Output
Section titled “4. Deterministic Output”Running the backup twice with no company changes should produce identical JSON except for backup metadata such as timestamps.
Backup Metadata
Section titled “Backup Metadata”The root object should begin with metadata similar to:
{ "archive_version": 1, "generated_at": "...", "company_id": "...", "company_name": "...", "realm_id": "...", "minor_version": "...", "sdk_version": "...", "entities": { }}Entity Collection Strategy
Section titled “Entity Collection Strategy”Create one function for each entity type.
Example
get_customers()
get_vendors()
get_accounts()
get_invoices()
get_payments()
...Each function should:
-
retrieve all pages
-
validate responses
-
return a list
Pagination
Section titled “Pagination”Implement one reusable pagination helper.
Responsibilities
-
page through all results
-
retry transient failures
-
stop only when complete
No entity should implement pagination itself.
Entity Coverage
Section titled “Entity Coverage”Attempt to archive every entity available to the authenticated company.
Typical examples include:
CompanyInfo
Preferences
Accounts
Customers
Vendors
Employees
Items
TaxCodes
TaxRates
Invoices
Payments
SalesReceipts
CreditMemos
RefundReceipts
Estimates
Bills
BillPayments
Purchases
Deposits
Transfers
JournalEntries
Checks
Credits
Attachables (metadata)
Classes
Departments / Locations
Budgets (if available)
Recurring transactions (if available)
Any additional supported entity exposed by the API.
Missing entity support should generate a warning, not terminate the backup.
Reports
Section titled “Reports”Reports are optional.
Since every report is derived from transactional data, reports are not required for disaster recovery.
However, consider archiving a small set of commonly used reports for validation purposes:
-
Balance Sheet
-
Profit & Loss
-
Trial Balance
These should be stored under:
validation_reportsError Handling
Section titled “Error Handling”Errors should be categorized.
Authentication
Authorization
Rate limiting
Temporary network
Validation
Unexpected API changes
Unexpected entities
Retry transient failures using exponential backoff.
Never silently discard data.
Logging
Section titled “Logging”Log
-
start
-
finish
-
entity counts
-
execution time
-
retries
-
warnings
-
failures
Prefer structured logging.
Validation
Section titled “Validation”After downloading:
Verify
-
entity counts
-
duplicate IDs
-
required identifiers
-
valid JSON
-
deterministic ordering
Emit warnings for anomalies.
JSON Layout
Section titled “JSON Layout”{ "metadata": { },
"entities": { "CompanyInfo": [ ],
"Customers": [ ],
"Vendors": [ ],
"Accounts": [ ],
"Invoices": [ ],
"Payments": [ ],
...
},
"validation_reports": { ... }}Performance
Section titled “Performance”Archive entity types independently.
Avoid holding unnecessary intermediate objects.
Use streaming serialization if the archive becomes very large.
Compress only if future testing shows a meaningful benefit. The primary artifact should remain an uncompressed JSON file for simplicity and maximum interoperability.
Security
Section titled “Security”Never archive
-
OAuth secrets
-
Refresh tokens
-
Client secrets
Never write credentials into logs.
Future Enhancements
Section titled “Future Enhancements”Potential future capabilities include:
-
Incremental backups
-
Daily scheduler
-
JSON schema validation
-
Restore utility
-
Archive comparison tool
-
Integrity hashing
-
Digital signatures
-
Optional compression
-
Optional encrypted archive
-
Automatic verification against previous backup
-
Multi-company support
Recommended JSON Viewers
Section titled “Recommended JSON Viewers”Do not build a custom HTML viewer.
Instead, use mature tooling.
Recommended viewers:
-
Visual Studio Code
-
Built-in JSON viewer
-
Folding
-
Search
-
Extensions such as JSON Crack for visualization if desired
-
-
Dadroit JSON Viewer
-
Excellent performance with very large JSON files
-
Tree navigation
-
Filtering
-
Fast search
-
-
jq
- Ideal for command-line inspection and scripting
-
Obsidian
- Useful once portions of the archive are transformed into Markdown for documentation or analysis, but not as the primary viewer.
Definition of Done
Section titled “Definition of Done”The project is complete when:
-
Authentication succeeds using the existing OAuth implementation.
-
Every supported entity can be downloaded from the target company.
-
Pagination is fully automatic.
-
The backup completes without manual intervention.
-
A single deterministic
QBO-archive.jsonfile is written toD:\FSS\Accounting\QuickBooks\. -
Validation completes successfully.
-
The archive can be opened and searched with standard JSON tools.
-
The process is robust enough to be scheduled and trusted as a long-term local archival solution.