The problem...

While working on my book/memoir, I wanted to find some meaningful distractions.


I had been using Scrivener for some time — which I would absolutely recommend, by the way — while constantly running ideas past ChatGPT in web chat. Back and forth, hashing out ideas, revising approaches, rereading things I had already written.


I’m by no means an author by trade, so having that feedback was incredibly useful. But chapter after chapter I kept thinking: this would be much better if ChatGPT could just be inside the application instead of making me tab over, reread everything, reconstruct the context, and add another multitasking nightmare to the pile.
At that point I started building my own writing application: Loom. Post on that one later.


As I started building Loom, I immediately ran into problem after problem trying to give an AI useful access to the application’s context.
Sometimes that context meant the page I was looking at as a whole. Sometimes it meant a paragraph or a passage. Maybe I wanted it to catch wording that was semantically wrong, a typo, or some abused punctuation. Other times I just wanted to ask:
“Hey, does this paragraph actually work within this outline?”


My original plan was to build an MCP server that exposed that context. But at the time, ChatGPT didn’t give me a practical way to use the custom/private MCP server I wanted, so I had to rethink how that context could be exposed.
That sent me down the rabbit hole of figuring out how MCP servers actually worked underneath.

Sucky solutions...

My first thought was simple enough: expose a port, build a REST interface into Loom, and let an AI client communicate with that.
Then came the obvious problem: my laptop isn’t a public web server.
So I tried ngrok — another service I’d absolutely recommend — and it worked almost immediately. I had a public HTTPS address pointing at the REST interface running inside Loom.


That solved reachability.
But reachability wasn’t really the whole problem.
An AI client still needed to know what was available at that endpoint, what it was allowed to call, and what each operation actually meant. I didn’t want to write a completely different set of hard-coded instructions for every application and every AI client.


That was the part of MCP that really stuck with me: the application describing its own capabilities instead of forcing the client to already know them.
If Loom could announce the context, actions, or commands it was willing to expose, then it wasn’t just ChatGPT that could potentially interact with it. Any compatible AI client could discover those capabilities.
That was the important part.

AI Proxy - Out of nowhere

Then the pieces started merging together.
What if the application didn’t have to be publicly exposed at all?
What if it connected outward, announced what it could do, and received a temporary public identity that an AI client could talk to?


I could borrow MCP’s capability-discovery pattern, some of the routing ideas behind a reverse proxy, and define a small contract between the public endpoint and the private application.


That became AI Proxy, and the contract became AI Proxy Protocol — APP.
The result was a way for clients like ChatGPT, Grok, Gemini, Muse, DeepSeek, or whatever came next to communicate with an application through only the capabilities that application deliberately exposed — without needing to stand up a dedicated MCP server for every combination.

How it works

At a high level, a downstream application connects outward to AP and creates a channel.


As part of that connection, it announces what it is capable of exposing: context, actions, commands, or whatever else the developer has deliberately made available.


AP gives that channel a public endpoint and a bootstrap prompt.
The “meat proxy” — me, in this case — hands that prompt to ChatGPT, Grok, Gemini, Muse, DeepSeek, or whatever client I happen to be using.
From there, AP sits in the middle.


The AI client can address the public channel, while the private application maintains the outbound relationship to AP, receives requests, executes only the capabilities it advertised, and sends responses back through the channel.
The important part is that the application itself never had to become a conventional publicly exposed service just so an AI client could reach it.

What it isn't...

AP isn’t intended to replace NGINX, MCP, a VPN, or some magical zero-trust security system.
Those things solve different problems.


MCP, for example, can still be one of the protocols exposed through AP. AP is concerned with a layer underneath that: how a private application establishes a presence, describes what it can do, and becomes reachable to an AI client without first turning itself into a traditional public service.

What will happen next?

There’s still plenty I want to do with it: stronger authentication, stable channel identities, better routing, multiple downstream applications, protocol adapters, and more.


But the basic idea already works.
Loom gave me a problem I thought I was solving for one writing application.
AP turned out to be the infrastructure I wished I’d had before I started.
Sometimes the distraction becomes the project.