This is the capstone. By the end you will have a repeatable deploy: edit HTML locally, run one command, see it live — and a verification pass that catches the mistakes this pipeline actually makes.
The approach is deliberately plain. Your source is a folder of .html files and a styles.css you can open in a browser from disk. No build step, no framework. The deploy script transforms them and writes them to page records.
How a page is stored
A page is a record of datatype page. What renders depends on which field is populated:
| Field | What it holds |
|---|---|
data.html | A complete HTML document. This is what you want |
data.screens[] | An array of { html } — a start/end wrapper pair |
data.content | The legacy virtual-DOM JSON, produced by the visual editor |
data.sections[] | Older section blocks — still rendered (as HTML) |
data.screens overrides data.html. If a page was ever saved in the visual editor it has screens, and your data.html edits will deploy successfully and change nothing on screen. This is the single most confusing failure in the whole pipeline, because every call returns 201.
When you take ownership of a page, clear the others in the same write:
{ "data.html": "…", "data.screens": [], "data.sections": [], "data.content": "" }The transforms
Three things have to happen between your source file and the record.
Inline the stylesheet
There is no static asset server for your CSS. <link rel="stylesheet" href="styles.css"> resolves against the route and 404s.
n = len(re.findall(r'<link rel="stylesheet" href="styles\.css"\s*/?>', html))
assert n == 1, f"{name}: expected 1 styles.css link, found {n}"
html = re.sub(r'<link rel="stylesheet" href="styles\.css"\s*/?>',
lambda m: "<style>\n" + css + "\n</style>", html, count=1)Match both <link …> and <link … />. A pattern that only handles one spelling silently does not substitute, and you deploy a page with no styling at all — which renders as a wall of unstyled text and looks like a far bigger problem than it is.
The assertion is what turns that into a failed deploy instead of a live incident. Assert the count before substituting, every time.
External stylesheets — Google Fonts and the like — are fine to leave as <link>s. It is only your own relative paths that break.
Rewrite internal links
Locally you link pricing.html. Deployed, the route is /pricing.
def rewrite(m):
page, tail = m.group(1), m.group(2) or '' # tail = #fragment or ?query
return 'href="' + ('/' if page == 'index' else '/' + page) + tail + '"'
html = re.sub(r'href="([a-z0-9\-]+)\.html([#?][^"]*)?"', rewrite, html)
leftover = re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html)
assert not leftover, f"{name}: unrewritten links {sorted(set(leftover))[:6]}"Handle the fragment. A pattern that requires the closing quote right after .html misses index.html#pricing, index.html#modules and every other anchored link — and those are exactly the links in your nav, on every page. They deploy verbatim and 404 for every visitor.
The leftover assertion catches the whole class. Run it on every page, every deploy.
Extract the metadata
title = re.search(r'<title>(.*?)</title>', html, re.S).group(1).strip()
md = re.search(r'<meta name="description" content="(.*?)"', html, re.S)
desc = md.group(1) if md else ""The deploy script
import json, os, re, sys, urllib.request, urllib.error
API = "https://appengine.appmint.io"
ORG = "your-org"
SRC = "/path/to/your/site"
UA = "Mozilla/5.0 (compatible; site-deploy)"
def req(method, path, body=None, token=None):
h = {"orgid": ORG, "User-Agent": UA, "Accept": "application/json"}
if body is not None: h["Content-Type"] = "application/json"
if token: h["Authorization"] = "Bearer " + token
r = urllib.request.Request(API + path, method=method, headers=h,
data=json.dumps(body).encode() if body is not None else None)
try:
with urllib.request.urlopen(r, timeout=90) as x:
return x.status, json.loads(x.read() or b"{}")
except urllib.error.HTTPError as e:
return e.code, e.read().decode("utf-8", "replace")[:400]
def build(name):
html = open(f"{SRC}/{name}.html").read()
css = open(f"{SRC}/styles.css").read()
n = len(re.findall(r'<link rel="stylesheet" href="styles\.css"\s*/?>', html))
assert n == 1, f"{name}: {n} styles.css links"
html = re.sub(r'<link rel="stylesheet" href="styles\.css"\s*/?>',
lambda m: "<style>\n" + css + "\n</style>", html, count=1)
def rw(m):
page, tail = m.group(1), m.group(2) or ''
return 'href="' + ('/' if page == 'index' else '/' + page) + tail + '"'
html = re.sub(r'href="([a-z0-9\-]+)\.html([#?][^"]*)?"', rw, html)
left = re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html)
assert not left, f"{name}: unrewritten links {sorted(set(left))[:6]}"
title = (re.search(r'<title>(.*?)</title>', html, re.S) or [None, name])[1].strip()
md = re.search(r'<meta name="description" content="(.*?)"', html, re.S)
return html, title, (md.group(1) if md else "")
def push(name, sk, token):
html, title, desc = build(name)
s, rec = req("GET", f"/repository/get/page/{sk}", token=token)
assert s == 200, (s, rec)
body = {"sk": sk, "version": rec.get("version", 0),
"data.html": html, "data.title": title, "data.description": desc}
if name == "index":
body.update({"data.screens": [], "data.sections": [], "data.content": ""})
s, d = req("POST", f"/repository/update-partial/page/{sk}", body, token=token)
print(f"{name:32s} {s}")
return s
ids = json.load(open("page-ids.json")) # {"index": "67dbad05…", "pricing": "6a63d834…"}
t = token()
for name in (sys.argv[1:] or ids.keys()):
push(name, ids[name], t)Keep page-ids.json in version control next to the source. Losing the id map means hunting records by name, and creating duplicates when you guess wrong.
Creating pages that do not exist yet
s, d = req("PUT", "/repository/create", {
"datatype": "page",
"name": name,
"data": {"name": name, "slug": name, "html": html,
"title": title, "description": desc},
}, token=t)
ids[name] = d["sk"]
json.dump(ids, open("page-ids.json", "w"), indent=2)Write the id back to the map immediately. A create that succeeds while the map is not saved leaves an orphan record, and the next run creates a second page with the same name.
Page JavaScript
Nothing in <script> tags will run — the parser discards script bodies. Page JavaScript belongs in the record's top-level style.javascript field, a sibling of data, which the renderer injects into the head as a real script.
Keep each page's script as a real .js file beside its HTML and send it up with the page:
js_path = f"{SRC}/{name}.js"
body = {"sk": sk, "version": rec.get("version", 0),
"data.html": html, "data.title": title, "data.description": desc}
if os.path.exists(js_path):
body["style"] = {"javascript": open(js_path).read()}Watch the nesting: style sits beside data, not inside it. data.style.javascript is read by nothing and fails silently.
See Run JavaScript on a hosted page for the full picture, including the attribute bootstrap for pipelines that can only write HTML.
Verify the deploy
A 201 means the record changed. It says nothing about what a visitor sees. Three checks, in order of what they catch.
Status codes are not a validity check
On this platform an unknown path returns 200 with the site's fallback body. curl -o /dev/null -w '%{http_code}' returns 200 for every URL you can invent, including /does-not-exist-xyz.
Check the <h1> instead:
def check(route):
with urllib.request.urlopen(
urllib.request.Request(BASE + route, headers={"User-Agent": UA})) as x:
html = x.read().decode("utf-8", "replace")
m = re.search(r"<h1[^>]*>(.*?)</h1>", html, re.S)
h1 = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", "", m.group(1))).strip() if m else "(no h1)"
return h1A route that renders the generic site name where you expected a page title is a 404 wearing a 200.
Crawl your own links
Every link, on every page, resolved:
targets = set()
for name in ids:
route = "/" if name == "index" else "/" + name
html = fetch(route)
assert not re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html), \
f"{route}: unrewritten .html links still live"
targets |= {l.split("#")[0].split("?")[0]
for l in re.findall(r'href="(/[^"]*)"', html)}
for t_ in sorted(targets):
print(t_, check(t_))The served HTML contains your markup more than once — the rendered DOM plus the framework's serialised payload, where < appears as <. A naive grep over the raw response counts each element two or three times. Match on the un-escaped copy, or parse the DOM in a browser, before concluding you have duplicates.
Look at it
Render the page and look. Curl greps are especially misleading here, because the serialised payload embeds the whole document — a title grep succeeds on a page whose body is blank.
Check, at least: the nav renders styled (catches a failed CSS inline), the footer is present, and any page script has run (dataset.bound on the element it binds to).
Things that bite
A helper that writes the file after a loop. If the loop asserts on a later item, the write never happens and earlier edits are lost silently. Write each file as you finish it, or make the whole pass atomic — not something in between.
Deploying a page you have not read. find returns a partial projection. Fetch with get before you edit, or fields you never saw will be dropped.
Nested objects in update-partial. {"data": {...}} replaces all of data. Dot-paths — data.html, data.title — change one key each.
A stale version. Always send the one from the get you just did.
Assuming the CDN has caught up. Custom domains lag the platform host by a cache cycle. Verify on <site>-<org>.site.appmint.app first; if the custom domain still shows the old page a minute later, it is a cache, not your deploy.
What the platform injects
The rendered page is not only your markup. The host adds its own chrome, and two consequences are worth knowing:
- Icon links in your
<head>are stripped. Favicons come from the site record — see Set a logo and favicon. - A cart drawer is present on every page, including a
Continue Shopping → /storelink. On a site with no storefront that route does not exist. It is hidden unless the drawer opens, and it comes from the host rather than your page, so it is not something a page edit can fix.
Checklist
-
styles.cssinlined, with an assertion on the link count - Internal links rewritten including fragments, with a leftover assertion
-
data.screens/sections/contentcleared on any page that had them - Page JavaScript in top-level
style.javascript(besidedata, not inside it) -
page-ids.jsoncommitted and written back on every create - Every route verified by
<h1>, not status code - At least one page rendered and looked at
Where to go next
- Client Integration → The server proxy — when a folder of HTML stops being enough
- The browser runtime — cart, products and drawers with no script
- Collect form submissions and Take bookings — the two things a marketing site nearly always needs