docs
/
Walkthroughs

Build and deploy a whole site

Take a folder of hand-written HTML to a live multi-page site — how pages are stored, a deploy script that is safe to re-run, the rewrites the platform expects, and how to verify a deploy without trusting status codes.

This is the capstone. By the end you will have a repeatable deploy: edit HTML locally, run one command, see it live — and a verification pass that catches the mistakes this pipeline actually makes.

The approach is deliberately plain. Your source is a folder of .html files and a styles.css you can open in a browser from disk. No build step, no framework. The deploy script transforms them and writes them to page records.

How a page is stored

A page is a record of datatype page. What renders depends on which field is populated:

FieldWhat it holds
data.htmlA complete HTML document. This is what you want
data.screens[]An array of { html } — a start/end wrapper pair
data.contentThe legacy virtual-DOM JSON, produced by the visual editor
data.sections[]Older section blocks — still rendered (as HTML)

data.screens overrides data.html. If a page was ever saved in the visual editor it has screens, and your data.html edits will deploy successfully and change nothing on screen. This is the single most confusing failure in the whole pipeline, because every call returns 201.

When you take ownership of a page, clear the others in the same write:

{ "data.html": "…", "data.screens": [], "data.sections": [], "data.content": "" }

The transforms

Three things have to happen between your source file and the record.

Inline the stylesheet

There is no static asset server for your CSS. <link rel="stylesheet" href="styles.css"> resolves against the route and 404s.

n = len(re.findall(r'<link rel="stylesheet" href="styles\.css"\s*/?>', html))
assert n == 1, f"{name}: expected 1 styles.css link, found {n}"
html = re.sub(r'<link rel="stylesheet" href="styles\.css"\s*/?>',
              lambda m: "<style>\n" + css + "\n</style>", html, count=1)

Match both <link …> and <link … />. A pattern that only handles one spelling silently does not substitute, and you deploy a page with no styling at all — which renders as a wall of unstyled text and looks like a far bigger problem than it is.

The assertion is what turns that into a failed deploy instead of a live incident. Assert the count before substituting, every time.

External stylesheets — Google Fonts and the like — are fine to leave as <link>s. It is only your own relative paths that break.

Locally you link pricing.html. Deployed, the route is /pricing.

def rewrite(m):
    page, tail = m.group(1), m.group(2) or ''     # tail = #fragment or ?query
    return 'href="' + ('/' if page == 'index' else '/' + page) + tail + '"'

html = re.sub(r'href="([a-z0-9\-]+)\.html([#?][^"]*)?"', rewrite, html)

leftover = re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html)
assert not leftover, f"{name}: unrewritten links {sorted(set(leftover))[:6]}"

Handle the fragment. A pattern that requires the closing quote right after .html misses index.html#pricing, index.html#modules and every other anchored link — and those are exactly the links in your nav, on every page. They deploy verbatim and 404 for every visitor.

The leftover assertion catches the whole class. Run it on every page, every deploy.

Extract the metadata

title = re.search(r'<title>(.*?)</title>', html, re.S).group(1).strip()
md = re.search(r'<meta name="description" content="(.*?)"', html, re.S)
desc = md.group(1) if md else ""

The deploy script

import json, os, re, sys, urllib.request, urllib.error

API = "https://appengine.appmint.io"
ORG = "your-org"
SRC = "/path/to/your/site"
UA  = "Mozilla/5.0 (compatible; site-deploy)"

def req(method, path, body=None, token=None):
    h = {"orgid": ORG, "User-Agent": UA, "Accept": "application/json"}
    if body is not None: h["Content-Type"] = "application/json"
    if token: h["Authorization"] = "Bearer " + token
    r = urllib.request.Request(API + path, method=method, headers=h,
        data=json.dumps(body).encode() if body is not None else None)
    try:
        with urllib.request.urlopen(r, timeout=90) as x:
            return x.status, json.loads(x.read() or b"{}")
    except urllib.error.HTTPError as e:
        return e.code, e.read().decode("utf-8", "replace")[:400]

def build(name):
    html = open(f"{SRC}/{name}.html").read()
    css  = open(f"{SRC}/styles.css").read()

    n = len(re.findall(r'<link rel="stylesheet" href="styles\.css"\s*/?>', html))
    assert n == 1, f"{name}: {n} styles.css links"
    html = re.sub(r'<link rel="stylesheet" href="styles\.css"\s*/?>',
                  lambda m: "<style>\n" + css + "\n</style>", html, count=1)

    def rw(m):
        page, tail = m.group(1), m.group(2) or ''
        return 'href="' + ('/' if page == 'index' else '/' + page) + tail + '"'
    html = re.sub(r'href="([a-z0-9\-]+)\.html([#?][^"]*)?"', rw, html)
    left = re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html)
    assert not left, f"{name}: unrewritten links {sorted(set(left))[:6]}"

    title = (re.search(r'<title>(.*?)</title>', html, re.S) or [None, name])[1].strip()
    md = re.search(r'<meta name="description" content="(.*?)"', html, re.S)
    return html, title, (md.group(1) if md else "")

def push(name, sk, token):
    html, title, desc = build(name)
    s, rec = req("GET", f"/repository/get/page/{sk}", token=token)
    assert s == 200, (s, rec)

    body = {"sk": sk, "version": rec.get("version", 0),
            "data.html": html, "data.title": title, "data.description": desc}
    if name == "index":
        body.update({"data.screens": [], "data.sections": [], "data.content": ""})

    s, d = req("POST", f"/repository/update-partial/page/{sk}", body, token=token)
    print(f"{name:32s} {s}")
    return s

ids = json.load(open("page-ids.json"))     # {"index": "67dbad05…", "pricing": "6a63d834…"}
t = token()
for name in (sys.argv[1:] or ids.keys()):
    push(name, ids[name], t)

Keep page-ids.json in version control next to the source. Losing the id map means hunting records by name, and creating duplicates when you guess wrong.

Creating pages that do not exist yet

s, d = req("PUT", "/repository/create", {
    "datatype": "page",
    "name": name,
    "data": {"name": name, "slug": name, "html": html,
             "title": title, "description": desc},
}, token=t)
ids[name] = d["sk"]
json.dump(ids, open("page-ids.json", "w"), indent=2)

Write the id back to the map immediately. A create that succeeds while the map is not saved leaves an orphan record, and the next run creates a second page with the same name.

Page JavaScript

Nothing in <script> tags will run — the parser discards script bodies. Page JavaScript belongs in the record's top-level style.javascript field, a sibling of data, which the renderer injects into the head as a real script.

Keep each page's script as a real .js file beside its HTML and send it up with the page:

js_path = f"{SRC}/{name}.js"
body = {"sk": sk, "version": rec.get("version", 0),
        "data.html": html, "data.title": title, "data.description": desc}
if os.path.exists(js_path):
    body["style"] = {"javascript": open(js_path).read()}

Watch the nesting: style sits beside data, not inside it. data.style.javascript is read by nothing and fails silently.

See Run JavaScript on a hosted page for the full picture, including the attribute bootstrap for pipelines that can only write HTML.

Verify the deploy

A 201 means the record changed. It says nothing about what a visitor sees. Three checks, in order of what they catch.

Status codes are not a validity check

On this platform an unknown path returns 200 with the site's fallback body. curl -o /dev/null -w '%{http_code}' returns 200 for every URL you can invent, including /does-not-exist-xyz.

Check the <h1> instead:

def check(route):
    with urllib.request.urlopen(
            urllib.request.Request(BASE + route, headers={"User-Agent": UA})) as x:
        html = x.read().decode("utf-8", "replace")
    m = re.search(r"<h1[^>]*>(.*?)</h1>", html, re.S)
    h1 = re.sub(r"\s+", " ", re.sub(r"<[^>]+>", "", m.group(1))).strip() if m else "(no h1)"
    return h1

A route that renders the generic site name where you expected a page title is a 404 wearing a 200.

Every link, on every page, resolved:

targets = set()
for name in ids:
    route = "/" if name == "index" else "/" + name
    html = fetch(route)
    assert not re.findall(r'href="(?!https?:)[^"]*\.html[^"]*"', html), \
        f"{route}: unrewritten .html links still live"
    targets |= {l.split("#")[0].split("?")[0]
                for l in re.findall(r'href="(/[^"]*)"', html)}

for t_ in sorted(targets):
    print(t_, check(t_))

The served HTML contains your markup more than once — the rendered DOM plus the framework's serialised payload, where < appears as <. A naive grep over the raw response counts each element two or three times. Match on the un-escaped copy, or parse the DOM in a browser, before concluding you have duplicates.

Look at it

Render the page and look. Curl greps are especially misleading here, because the serialised payload embeds the whole document — a title grep succeeds on a page whose body is blank.

Check, at least: the nav renders styled (catches a failed CSS inline), the footer is present, and any page script has run (dataset.bound on the element it binds to).

Things that bite

A helper that writes the file after a loop. If the loop asserts on a later item, the write never happens and earlier edits are lost silently. Write each file as you finish it, or make the whole pass atomic — not something in between.

Deploying a page you have not read. find returns a partial projection. Fetch with get before you edit, or fields you never saw will be dropped.

Nested objects in update-partial. {"data": {...}} replaces all of data. Dot-paths — data.html, data.title — change one key each.

A stale version. Always send the one from the get you just did.

Assuming the CDN has caught up. Custom domains lag the platform host by a cache cycle. Verify on <site>-<org>.site.appmint.app first; if the custom domain still shows the old page a minute later, it is a cache, not your deploy.

What the platform injects

The rendered page is not only your markup. The host adds its own chrome, and two consequences are worth knowing:

  • Icon links in your <head> are stripped. Favicons come from the site record — see Set a logo and favicon.
  • A cart drawer is present on every page, including a Continue Shopping → /store link. On a site with no storefront that route does not exist. It is hidden unless the drawer opens, and it comes from the host rather than your page, so it is not something a page edit can fix.

Checklist

  • styles.css inlined, with an assertion on the link count
  • Internal links rewritten including fragments, with a leftover assertion
  • data.screens/sections/content cleared on any page that had them
  • Page JavaScript in top-level style.javascript (beside data, not inside it)
  • page-ids.json committed and written back on every create
  • Every route verified by <h1>, not status code
  • At least one page rendered and looked at

Where to go next