Skip to content

QuickBooks Online Local Backup (QBO-Archive)

Section titled “QuickBooks Online Local Backup (QBO-Archive)”

Create a deterministic, local backup of an entire QuickBooks Online company using the Intuit Developer API.

The backup is read-only and intended solely for disaster recovery, historical preservation, migration assistance, and offline inspection.

The output is a single JSON file.

No attempt is made to restore the data automatically. A future restore tool may consume this JSON.


The backup system should be:

  • Complete

  • Deterministic

  • Repeatable

  • Idempotent

  • Easy to validate

  • Human inspectable

  • AI inspectable

  • Independent of Intuit

The system should produce exactly one file:

D:\FSS\Accounting\QuickBooks\QBO-archive.json

The file is overwritten each execution.

Historical versions are retained by Kopia.


Python

Package manager

  • uv

Libraries

  • httpx

  • pydantic

  • orjson

  • loguru

Reuse the existing OAuth implementation already used by the transaction-processing project.

No third-party QuickBooks wrappers unless they provide a clear long-term maintenance advantage.


qbo-archive/
pyproject.toml
src/
archive.py
api.py
auth.py
config.py
models.py
pagination.py
serializer.py
main.py

The backup should retrieve every supported entity available through the QBO REST API that the authenticated company exposes.

Never back up only reports.

Always back up the underlying source objects.


Do not normalize fields.

Do not rename fields.

Do not remove unknown properties.

The backup should preserve the API response exactly whenever practical.

If additional metadata is added by the backup system, place it outside the raw API objects.


Objects must always be sorted before writing.

Recommended ordering:

  • entity type

  • primary identifier

  • creation date

  • update timestamp

Stable ordering makes Git diffs and backup comparisons much easier.


Running the backup twice with no company changes should produce identical JSON except for backup metadata such as timestamps.


The root object should begin with metadata similar to:

{
"archive_version": 1,
"generated_at": "...",
"company_id": "...",
"company_name": "...",
"realm_id": "...",
"minor_version": "...",
"sdk_version": "...",
"entities": { }
}

Create one function for each entity type.

Example

get_customers()
get_vendors()
get_accounts()
get_invoices()
get_payments()
...

Each function should:

  • retrieve all pages

  • validate responses

  • return a list


Implement one reusable pagination helper.

Responsibilities

  • page through all results

  • retry transient failures

  • stop only when complete

No entity should implement pagination itself.


Attempt to archive every entity available to the authenticated company.

Typical examples include:

CompanyInfo

Preferences

Accounts

Customers

Vendors

Employees

Items

TaxCodes

TaxRates

Invoices

Payments

SalesReceipts

CreditMemos

RefundReceipts

Estimates

Bills

BillPayments

Purchases

Deposits

Transfers

JournalEntries

Checks

Credits

Attachables (metadata)

Classes

Departments / Locations

Budgets (if available)

Recurring transactions (if available)

Any additional supported entity exposed by the API.

Missing entity support should generate a warning, not terminate the backup.


Reports are optional.

Since every report is derived from transactional data, reports are not required for disaster recovery.

However, consider archiving a small set of commonly used reports for validation purposes:

  • Balance Sheet

  • Profit & Loss

  • Trial Balance

These should be stored under:

validation_reports

Errors should be categorized.

Authentication

Authorization

Rate limiting

Temporary network

Validation

Unexpected API changes

Unexpected entities

Retry transient failures using exponential backoff.

Never silently discard data.


Log

  • start

  • finish

  • entity counts

  • execution time

  • retries

  • warnings

  • failures

Prefer structured logging.


After downloading:

Verify

  • entity counts

  • duplicate IDs

  • required identifiers

  • valid JSON

  • deterministic ordering

Emit warnings for anomalies.


{
"metadata": { },
"entities":
{
"CompanyInfo": [ ],
"Customers": [ ],
"Vendors": [ ],
"Accounts": [ ],
"Invoices": [ ],
"Payments": [ ],
...
},
"validation_reports":
{
...
}
}

Archive entity types independently.

Avoid holding unnecessary intermediate objects.

Use streaming serialization if the archive becomes very large.

Compress only if future testing shows a meaningful benefit. The primary artifact should remain an uncompressed JSON file for simplicity and maximum interoperability.


Never archive

  • OAuth secrets

  • Refresh tokens

  • Client secrets

Never write credentials into logs.


Potential future capabilities include:

  • Incremental backups

  • Daily scheduler

  • JSON schema validation

  • Restore utility

  • Archive comparison tool

  • Integrity hashing

  • Digital signatures

  • Optional compression

  • Optional encrypted archive

  • Automatic verification against previous backup

  • Multi-company support


Do not build a custom HTML viewer.

Instead, use mature tooling.

Recommended viewers:

  1. Visual Studio Code

    • Built-in JSON viewer

    • Folding

    • Search

    • Extensions such as JSON Crack for visualization if desired

  2. Dadroit JSON Viewer

    • Excellent performance with very large JSON files

    • Tree navigation

    • Filtering

    • Fast search

  3. jq

    • Ideal for command-line inspection and scripting
  4. Obsidian

    • Useful once portions of the archive are transformed into Markdown for documentation or analysis, but not as the primary viewer.

The project is complete when:

  • Authentication succeeds using the existing OAuth implementation.

  • Every supported entity can be downloaded from the target company.

  • Pagination is fully automatic.

  • The backup completes without manual intervention.

  • A single deterministic QBO-archive.json file is written to D:\FSS\Accounting\QuickBooks\.

  • Validation completes successfully.

  • The archive can be opened and searched with standard JSON tools.

  • The process is robust enough to be scheduled and trusted as a long-term local archival solution.