You have a JSON file of products — scraped, exported from another platform, or written by hand — and you need them in the storefront.
The mechanics are simple. What costs time is the handful of places the record shape is not what you would guess, and the fact that a half-finished load is much worse than no load at all.
The record
payload = {
"name": name[:100],
"datatype": "sf_product",
"isNew": True,
"author": AUTH_EMAIL,
"data": {
"name": name,
"slug": slug,
"sku": sku,
"price": product.get("price", 0),
"description": product.get("description", ""),
"status": "available",
"available": True,
},
"post": {
"categories": product.get("categories", []),
"tags": product.get("tags", []),
},
}
resp = requests.put(f"{API_BASE}/repository/create", headers=headers, json=payload)Four things in there are not obvious.
isNew: True is required
Leave it out and every create fails:
{ "statusCode": 400,
"error": "create => Not a new metrics, please use update or set the new property" }The flag is isNew. The message says "the new property", which reads as a field called new — it is not, and chasing that wording is a dead end.
Categories and tags go in post, not data
"post": { "categories": [...], "tags": [...] }Put them in data and they are stored — as inert fields nothing reads. Category listings and tag filters resolve against post, so the products load successfully and then appear in no category at all.
name is truncated at the top level
name[:100]. The top-level name is the record's label and has a length limit; data.name carries the real product title. Send a 300-character name and the create fails for a reason the error will not make obvious.
Images are objects, and there are two fields
data["images"] = uploaded # [{url, path}, …]
data["image"] = uploaded[0]["url"] # a thumbnail stringimages is an array of FileInfo-shaped objects. image is a plain string for the thumbnail. Populating only images gives you a catalogue where every tile is blank.
Images
Upload each file first, then attach what comes back:
def upload_image(token, image_path, slug):
location = f"{ORG_ID}/products/{slug}" # FOLDER only
with open(image_path, "rb") as f:
resp = requests.post(
f"{API_BASE}/repository/file/upload",
headers={"Authorization": f"Bearer {token}", "orgid": ORG_ID},
files={"file": (os.path.basename(image_path), f)},
data={"location": location},
timeout=120,
)
result = resp.json()
path = result.get("path", "")
url = result.get("signedUrl") or result.get("url") or f"{CDN_BASE}/{path}"
return {"url": url, "path": path}location is the folder, not the file path. The endpoint appends the filename itself. Pass products/chair/chair.jpg and you get products/chair/chair.jpg/chair.jpg.
Two more things about uploads at volume:
- Raise the timeout. 120 seconds, not the default. A large image on a slow link will otherwise fail partway through a long run.
- Uploads are served as
application/octet-stream. Browsers sniff raster images, so JPEG and PNG render fine. SVG does not — a browser refuses to render an SVG served that way. If your catalogue has SVG imagery, convert it or inline it.
Make it re-runnable
This matters more than anything else on this page. A bulk load will fail partway through — a timeout, a bad row, an expired token — and if re-running it duplicates everything you now have a worse problem than you started with.
Fetch what exists first, and skip it:
def get_existing_skus(token):
skus, page = set(), 1
while True:
resp = requests.post(
f"{API_BASE}/repository/find/sf_product",
headers=headers,
json={"query": {}, "options": {"pageSize": 100, "page": page}},
timeout=30,
)
data = resp.json()
for item in data.get("data", []):
sku = item.get("data", {}).get("sku")
if sku:
skus.add(sku)
if not data.get("hasNext"):
break
page += 1
return skusThen the load is idempotent:
existing = get_existing_skus(token)
for p in products:
if p["sku"] in existing:
print(f"skip {p['sku']}")
continue
create_product(token, p, images)Page with hasNext, not by counting. A loop that stops when it has seen total records will loop forever if anything is written while it runs.
Deduplicate on a business key — the SKU — not on the record id. You do not have ids for rows you have not created yet, which is exactly the case you are guarding against.
Take a range
python3 upload-products.py 0 50 # products 0–49Loading fifty at a time turns a long run into a series of short ones you can inspect between. On a first run against a live org this is the difference between one bad field and four hundred records carrying it.
Order of operations
- Validate the JSON offline. Every row has a SKU, a name, a numeric price. Fail the whole run on a bad row rather than discovering it at record 300.
- Load five. Look at them in the storefront — tile, detail page, category, search.
- Fix what is wrong and delete those five.
- Load the rest in ranges.
- Reconcile — count what is in the org against what is in the file.
Step 2 is the one people skip. A field in the wrong place is invisible in the API response and obvious the moment you look at a product tile.
Updating, not creating
For a re-load that must update existing rows, find by SKU and patch with dot-paths:
s, d = req("POST", "/repository/query/sf_product", {
"where": [{"field": "data.sku", "operator": "eq", "value": sku}]
}, token=t)
rec = d["data"][0]
req("POST", f"/repository/update-partial/sf_product/{rec['sk']}", {
"sk": rec["sk"], "version": rec["version"],
"data.price": new_price,
}, token=t)Never send {"data": {...}} in an update. It replaces the entire data object — images, parcel, attributes and everything else you did not include are gone. Dot-paths change one key each. This has wiped prices and image lists in a real migration; the recovery was a find snapshot taken beforehand, which is the only reason it was recoverable.
Snapshot before any bulk write:
curl -s "$APPMINT_HOST/repository/find/sf_product?perPage=1000" \
-H "orgid: $APPMINT_ORG" -H "Authorization: Bearer $TOKEN" \
> backup-$(date +%Y%m%d-%H%M).jsonMigrating from another platform
Moving a real store — 36.7K records off WooCommerce, in the case this is drawn from — the shape holds but two things change:
Go through an intermediate format. Source database → SQL dump → JSONL → uploads. Each step is inspectable and re-runnable. Going straight from a live database to the API means any failure restarts the whole thing.
Customers cannot be created like products. A customer record is an identity with credentials and sessions behind it. They come in through the signup path, not repository/create.
Order matters: products, then customers, then orders — because orders reference both.
Media is a phase of its own. Get the catalogue in and correct with a placeholder image first; a migration that blocks on 20,000 image transfers stalls before anyone can check whether the product data is right.
Checklist
-
isNew: Trueon every create - Categories and tags in
post, notdata - Top-level
nametruncated to 100 - Both
images(array of objects) andimage(string) set - Upload
locationis a folder - Existing SKUs fetched and skipped — the load is re-runnable
- Paged with
hasNext - Five loaded and eyeballed before the rest
- Snapshot taken before any update pass
- Updates use dot-paths
Next: HTML to live pages — the same discipline, applied to a whole site.