DHDGUI-HyperMem Docs
Copy for LLMGet a token
Memory infrastructure for agents

DGUI-HyperMem documentation

Connect an AI client to a scoped hybrid memory service with MCP, retrieve durable context with vector and keyword search, and manage the same data through REST.

Hosted endpoint available MCP over Streamable HTTP Cloudflare self-hostable
For users

Start with a token

Request access, store a memory, search it back, and remove it when you no longer want it retained.

Connect in five steps
For developers

Use MCP or REST

Use the authenticated MCP endpoint for agents or the REST mirror for scripts, services, and diagnostics.

Inspect the interface
For operators

Run your own

Deploy the Worker with D1, Vectorize, Workers AI, and secrets managed by your Cloudflare account.

Self-host the service
Hosted endpoint: https://dgui-hypermem.ctaxnagomi.workers.dev. The MCP endpoint is /mcp; the service status endpoint is /health.

Establish your connection

Complete this sequence once per client or application. Keep the bearer token in an environment variable or a client secret prompt, not in source code.

Request access

Open the hosted service, accept the terms, and request a token. Store the returned token as DGUI_HYPERMEM_TOKEN.

Expose the token to your shell

Use a secret manager or a local environment variable. PowerShell: $env:DGUI_HYPERMEM_TOKEN = "YOUR_TOKEN". Bash: export DGUI_HYPERMEM_TOKEN="YOUR_TOKEN".

Add the remote MCP server

Copy the block for your client from Client guides. Do not replace the URL with the landing-page URL; the server URL must end in /mcp.

Verify authentication

Call GET /api/verify-token with the bearer token. A valid CRM token returns valid: true and its effective quota.

Run a memory smoke test

Use the MCP add tool, then search with the same scope. Remove the test record with forget and its returned ID.

REST verification
curl "https://dgui-hypermem.ctaxnagomi.workers.dev/api/verify-token" \
  -H "Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}"
Agent instruction: “Before starting a project, search DGUI-HyperMem for relevant decisions and preferences. After a durable decision is made, add a concise memory with a stable project scope and useful tags. Never store credentials or private keys.”

User guide

DGUI-HyperMem is useful when information should survive a chat, coding session, client restart, or model switch. Use concise statements that will still make sense later.

RecallSearch before asking the user to repeat prior context.
DecideRecord durable choices, constraints, and preferences.
VerifyCheck the returned ID, type, tags, and source.
PruneForget stale or incorrect records by ID.

add

Stores a memory, assigns a JEV type and salience, checks nearby memories for contradictions, and may supersede an older record.

search

Combines Vectorize similarity, D1 FTS5 keyword search, reciprocal-rank fusion, and optional JEV reranking.

list

Browses active, superseded, or deleted records in one scope, newest first, with type and status filters.

profile

Summarizes counts by type, popular tags, average salience, and recent activity for a scope.

forget

Deletes exact IDs, the closest matches to a query, or an entire scope. Prefer IDs for routine cleanup.

help

Returns the server’s built-in tool summary and current JEV provider behavior.

Writing useful memories

  • Use stable scopes such as project:hypermem, user:default, or agent:research.
  • Write one durable fact or decision per memory instead of a transcript.
  • Include constraints and rationale when a future agent needs to understand why.
  • Add tags for technologies, subsystems, environments, and workflows.
  • Set source to a useful provenance label without embedding secrets or personal data.
Privacy: memories are retained until explicitly forgotten or the relevant scope is cleared. JEV analysis can queue decision examples for dataset synchronization when an AI provider and dataset export are configured. Do not store passwords, API keys, private keys, payment data, or unnecessary personal information.

Developer guide

The Worker exposes a stateless MCP server and a small REST mirror. Both operate on the same D1 records and Vectorize index.

Retrieval pipeline

1. Embed@cf/baai/bge-base-en-v1.5 produces 768 dimensions.
2. RetrieveVectorize ANN and D1 FTS5 BM25 return independent candidates.
3. FuseReciprocal rank fusion merges both candidate lists.
4. RerankJEV reorders results when its configured provider is available.

Memory behavior

add accepts every non-empty memory that passes the content filter. JEV assigns type, salience, confidence, and durability metadata; a durability score is not an automatic deletion rule. A nearby active record can be marked superseded when the contradiction score reaches the configured threshold.

Base URLs

PurposeURLAuthentication
MCPhttps://dgui-hypermem.ctaxnagomi.workers.dev/mcpBearer token
RESThttps://dgui-hypermem.ctaxnagomi.workers.dev/api/*Bearer token
Statushttps://dgui-hypermem.ctaxnagomi.workers.dev/healthPublic

Client guides

Each block uses the hosted endpoint and the DGUI_HYPERMEM_TOKEN environment variable. For self-hosting, replace the URL with your Worker URL and keep the /mcp suffix.

IDE

VS Code

Uses top-level servers and secret inputs with ${input:dgui_hypermem_token} interpolation.

Copy VS Code config
Terminal agent

Claude Code

Registers a remote HTTP MCP server and sends the token through an authorization header.

Copy Claude Code command
IDE

JetBrains

Stores the remote server in a project or user MCP JSON file with an authorization header.

Copy JetBrains config
opencode: current opencode configuration places remote servers directly under mcp, uses type: "remote", reads secrets with {env:NAME}, and supports bearer headers. Copy the opencode block.

Authentication

The /mcp endpoint accepts two credential forms. OAuth 2.1 is the recommended path for MCP clients: the client discovers the endpoints, registers itself, and the user signs in through a browser consent flow — no token to copy or store. A bearer token remains supported for scripts and for clients with no OAuth support.

OAuth 2.1

The Worker is its own authorization server, so there is no third-party identity provider and no per-user vendor cost. Clients discover the configuration, then run an authorization-code flow with PKCE:

Discovery documents
GET /.well-known/oauth-protected-resource
GET /.well-known/oauth-authorization-server

POST /register     dynamic client registration (RFC 7591)
GET  /authorize    consent + sign-in
POST /token        code exchange and refresh
POST /revoke       token revocation (RFC 7009)

PKCE with S256 is mandatory, and authorization codes are single-use with a ten-minute lifetime. Access tokens last one hour and are rotated on refresh. All three token classes — authorization codes, access tokens, and refresh tokens — are stored only as SHA-256 digests, so a database leak does not yield usable credentials. Signing in uses the same email and passkey as the token page, so an OAuth grant resolves to the same account, quota, and usage history as a pasted token. Disabling an account immediately invalidates every OAuth session derived from it.

Most clients need nothing beyond the server URL — leave Authorization unset and let the client run the flow. When a request arrives unauthenticated, the endpoint answers 401 with a WWW-Authenticate header pointing at the resource metadata, which is how a standards-compliant client knows to begin.

Bearer tokens

Hosted requests also accept an active CRM token in the Authorization header. The Worker accepts X-API-Key and a token query parameter for compatibility, but bearer headers are preferred because query strings are commonly retained in logs and proxies.

Preferred authorization header
Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}
  • GET /api/verify-token validates status, terms acceptance, and quota state.
  • GET /api/check-quota returns plan, usage, remaining allowance, and reset time.
  • Do not paste a token into a prompt, URL, issue, screenshot, or tracked configuration file.
  • Self-hosted deployments must set MCP_TOKEN as a Worker secret. Authorization fails closed when it is unset, so an unset secret denies every request rather than admitting them.

Memory tools

ToolRequired inputOptional inputResult
addcontentscope, tags, source, check_contradictionsID, type, salience, durability, provider, superseded IDs
searchqueryscope, limit, type, durable_only, use_jevHybrid-scored memory results
listNonescope, limit, type, statusNewest memories in one scope
profileNonescopeCounts, tags, salience, recent activity
forgetOne of ids, query, or allscopeDeleted count and IDs
helpNoneNoneBuilt-in service summary
sync_jev_datasetNonelimit, 1–500Dataset upload result
jev_queue_statsNoneNonePending, uploaded, and failed queue counts
Deletion warning: forget with all: true clears the selected scope. Confirm the scope before invoking it and prefer exact IDs when correcting one record.

REST API

All routes below are under /api. Except for the public landing workflow routes, send the token as a bearer header. Mutation and search routes use POST with a JSON body.

RouteMethodBody or queryPurpose
/addPOSTcontent, scope, tags, sourceAdd or update a scoped memory
/searchPOSTquery, scope, limit, type, durable_onlyHybrid search
/listPOSTscope, limit, type, statusList memories
/profilePOSTscopeSummarize a scope
/forgetPOSTids, query, scope, or allDelete memories
/verify-tokenGETBearer tokenValidate a token
/check-quotaGETBearer tokenInspect effective usage, trial and wallet
/billing-summaryPOSTemail, passkeyPlan, trial and wallet before you hold a token
/start-trialPOSTemail, passkeyClaim the one-time 15-day Pro trial
/create-checkout-sessionPOSTplan, emailStart a subscription checkout
/buy-creditsPOSTpack, emailTop up pay-as-you-go credit
/sync_jevPOSTlimitFlush queued JEV examples
/jev_queue_statsGETOptional scope queryInspect dataset queue state
Add and search over REST
curl -X POST "https://dgui-hypermem.ctaxnagomi.workers.dev/api/add" \
  -H "Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{"content":"The project uses Cloudflare Workers.","scope":"project:example","tags":["cloudflare","workers"]}'

curl -X POST "https://dgui-hypermem.ctaxnagomi.workers.dev/api/search" \
  -H "Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{"query":"What platform does the project use?","scope":"project:example","limit":5}'

Scopes and tenancy

A scope is a logical namespace attached to every memory operation. Use stable names so the same client can recall the same context later.

Recommended patterns

project:repository-name
user:stable-id
agent:role-name
team:team-name

Default behavior

If a request omits scope, the Worker uses DEFAULT_SCOPE, then default.

Important: scopes organize data but are not authorization boundaries in the current implementation. A valid token can address any scope string. Use separate self-hosted deployments or add application-level authorization when scopes must isolate tenants.

Quotas

Usage is tracked per token against the plan allowance on a 30-day reset window: 2,000 requests for Free, 8,500 for Median, 15,000 for Pro, and 25,000 for Enterprise. Once the plan allowance is spent, requests continue against any prepaid pay-as-you-go balance before the request is refused.

Check effective quota
curl "https://dgui-hypermem.ctaxnagomi.workers.dev/api/check-quota" \
  -H "Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}"

The response includes plan, quota_monthly, requests_used, requests_remaining, resets_at, 30-day usage, and yearly usage. Use the response as the source of truth for an account because operator overrides take precedence over plan defaults.

Pipeline

Every write follows one path: the request is analyzed by the JEV reasoning layer, written to D1 and Vectorize, and — when a provider and a dataset export are both configured — queued as a training row that is later flushed to Hugging Face. Reads walk the same path in reverse through hybrid recall. Where the queued rows land depends on which dataset is bound, and the reasoning layer can be parked in standby without taking memory offline.

1. Ingestadd stores content, scope, tags, and source.
2. AnalyzeJEV assigns type, durability, and salience.
3. IndexD1 row plus a Vectorize embedding for recall.
4. QueueThe decision row enters the jev_examples queue.
5. Flushsync_jev_dataset or the hourly cron appends JSONL.
6. Recallsearch fuses both candidate lists, then re-ranks.

Private data

Point HF_DATASET at a repository you own and supply an HF_TOKEN that can write to it. Every flush then appends to train.jsonl and rewrites metadata.json in that repository instead of the public one. /api/setup-dataset validates a token against a repository and records the setup so an operator can confirm the binding before the first sync; the effective sync target always remains the HF_DATASET binding.

Two caveats before you rely on a private dataset. The Worker only auto-creates a repository when one does not exist yet, and it creates it as public — so create the private repository yourself first. And /api/setup-dataset records a short token prefix in the CRM audit log, so treat that log as sensitive.

Public data

With no override, rows go to the public ctaxnagomi/DGUI_HYPERMEM-JEV dataset, which is created on first flush if it is missing. Each row keeps its use_case, instruct_type, the exact questions sent, the answers returned, and the provider and model that produced them, so the corpus is usable as-is for few-shot or fine-tuning work.

Read the public corpus
from datasets import load_dataset

ds = load_dataset("ctaxnagomi/DGUI_HYPERMEM-JEV", split="train")

for row in ds.stream():
    print(row["use_case"], row["instruct_type"], row["provider"])

Training module

The queue is the training module. Every JEV decision becomes one typed instruction row in jev_examples with status pending, uploaded, or error, flushed in batches of 200 by the sync_jev_dataset tool or by the scheduled cron. Three use cases are recorded: analyze (type, durability, salience for a new memory), rerank (which candidates actually answer a query), and supersede (whether an incoming memory replaces a stale one).

ControlWhereEffect
jev_queue_statsMCP toolPending, uploaded, and error counts, a per-use-case breakdown, the last flush time, and the target repository.
sync_jev_datasetMCP toolFlushes up to limit pending rows immediately instead of waiting for the cron.
/api/sync_jevRESTThe same flush over HTTP for operators and scheduled jobs.
/api/admin/toggle-trainAdminFlips the train_with_all flag stored on a token.
Consent gap to be aware of. train_with_all is recorded per account and toggled from the admin console, but the current write path does not check it before queueing a row. Until that gate lands, treat dataset export as opt-in by leaving HF_TOKEN unset, and do not rely on the flag alone to keep an account's decisions out of the corpus.

Idle and standby mode

Standby is the non-serving state. Rather than spending capacity on requests, the deployment uses its idle time to let the JEV layer train on the decisions it has already accumulated and to assemble a new module on the corpus dedicated to prognostic skill. Standby changes what the system does with spare capacity, not what it will answer when a client does call — memory service stays available throughout.

Proposed capability. Idle/standby training and the prognostic corpus module are not implemented in the current Worker. What exists today is the decision queue and the flush described above, which is the substrate this mode would build on. Treat the remainder of this section as design intent rather than current behaviour.

What exists today

The queue already behaves as an append-only corpus of the system’s own reasoning, so the raw training signal for a prognostic module is accumulating whether or not that module exists. Three use cases are recorded today — analyze, rerank, and supersede — each flushed hourly or on demand, each carrying the questions asked and the answers returned. A prognostic module adds a fourth rather than replacing anything.

The prognostic module

The intended module is a prognose use case. Where the existing use cases score a memory, a prognosis scores an expectation: given the current state and recent decision history, what is likely to be needed or to happen next. A row would carry the question block, the expectation, its confidence, and — the part that makes it trainable — the observed outcome once it is known, so each expectation becomes a labelled example of being right or wrong.

That outcome field is what separates a prognostic corpus from a heuristic. Expectations recorded before the fact and scored after it are a supervised signal, so accuracy improves as usage accumulates instead of depending on prompt wording. It is also the honest way to expose uncertainty: an expectation with no outcome yet is pending, not correct.

Real-time expectation

The feature this unlocks is real-time expectation: as an agent works, the layer maintains a short, ranked set of expectations about what comes next rather than waiting to be asked. In practice that means surfacing the memories and constraints most likely to be needed instead of the merely textually similar ones, flagging when the current direction contradicts a stored expectation, and recording what was expected at the moment a decision is made so the next cycle has a baseline to score against. It fits the existing primitives well, because a prognosis is a graded judgement rather than a hard classification — the same shape the current layer already uses for durability and salience.

Designing for it

  • Keep the queue append-only and version the use_case field, so adding a module never rewrites rows already exported.
  • Keep outcome labelling separate from generation, so an unlabelled expectation is visibly pending rather than silently counted as correct.
  • Expect inference cost to scale with active scopes, so keep a standby pass bounded, resumable, and interruptible by a foreground request.
  • Honour train_with_all: prognosis derived from an account’s decisions should follow the same consent rule as any other exported row, and the gap noted above applies to this module too.

Pausing the reasoning layer

Separately from standby, JEV_MODE controls whether the reasoning layer runs at all. Setting it to off (or none or false) parks analysis without stopping the memory service: add and search still store and retrieve, embeddings are still written, and hybrid recall still fuses vector and keyword candidates — but every analysis reports provider: "off", memories fall back to a neutral classification, re-ranking is skipped, and nothing is appended to the training queue.

JEV_MODEResolved providerBehaviour
auto (default)typesafe when TYPESAFE_API_KEY is set, otherwise workers-aiFull analysis, re-ranking, and queueing.
typesafetypesafe, falling back to workers-ai without a keySame, forced toward the TypeSafe backend.
workers-aiworkers-aiSame, forced onto the Workers AI fallback.
offoffAnalysis paused: neutral classification, no re-ranking, no queueing, no provider calls. Memory service unaffected.
Two different meanings of “off”. Pausing the reasoning layer is a server-side binding and applies to the whole deployment; there is no per-token switch for it. It is also unrelated to idle/standby mode, which is about training rather than disabling analysis. Staff clock-in and clock-out (/api/admin/clock) is a separate human timekeeping feature and affects neither.

Monetize

The hosted service already meters usage per token and exposes the billing plumbing, so monetizing it is a matter of pricing, checkout, and packaging rather than new instrumentation. Everything below runs on the current Worker.

Plans and quotas

Each token carries a plan, counted per request on a 30-day reset window. Plan allowances are 2,000 for Free, 8,500 for Median, 15,000 for Pro, and 25,000 for Enterprise. A quota_override on the token, when set, always wins over the plan — that is how a custom enterprise allowance is granted without a code change. Leaving it NULL is what lets the plan ladder take effect.

PlanDefault monthly requestsTypical use
Free2,000Evaluation and a single small project.
Median8,500One active agent or a small team.
Pro15,000Multi-agent setups and CI integrations.
Enterprise25,000Self-hosted or a custom override.
Pay as you gounlimitedBilled per request once the plan allowance is spent.

/api/check-quota returns plan, quota_monthly, requests_used, requests_remaining, resets_at, 30-day and yearly totals, and a payg block with the wallet balance, what it buys at the current rate, and lifetime spend. When the plan allowance is spent it also returns exhausted: true and an upgrade block naming each option and its price, so a client can present the choices rather than a bare refusal. Use that response as the billing source of truth rather than the plan name, since operators can override a plan's allowance. Pay-as-you-go is charged at $0.002 per request, debited from the wallet only after the plan allowance is exhausted, and the debit is a guarded conditional update so concurrent requests cannot overdraw the balance.

Checkout and upgrades

Upgrades run through Stripe. /api/create-checkout-session starts a Checkout session and /api/stripe-webhook applies the result to the token — moving plan, status, and quota_monthly once payment settles. The upgrade page is served at /pay, with /payment and /upgrade as aliases. /api/admin/update-quota handles manual plan changes and overrides for accounts that cannot use card checkout.

Billing prerequisite. Checkout needs STRIPE_SECRET_KEY, and the webhook only settles orders when STRIPE_WEBHOOK_SECRET is also configured. With the webhook secret unset, a completed payment is never confirmed into the token, so verify both are present before advertising a paid tier.

Enterprise and self-hosted revenue

/api/enterprise-inquiry captures name, email, company, and message as an inbound lead. For larger accounts the stronger motion is a dedicated deployment: the Self-host section keeps D1 records, Vectorize vectors, provider usage, and any Hugging Face export inside the customer's own Cloudflare account, which is usually easier to sell than a shared hosted plan. Keep an Enterprise plan with a raised quota_monthly as the billing record for those deployments.

Packaging the offer

  • Lead with the agent loop, not the database. Clients configure one URL and immediately gain add, search, list, profile, and forget.
  • Sell scopes as workspaces. A scope per project, user, or agent is the natural unit to meter and the natural unit to export.
  • Charge for the reasoning layer. The JEV provider and the dataset flush are the differentiated parts; keep them on paid tiers and leave the free tier on the Workers AI fallback.
  • Let the corpus be the marketing. The public dataset shows real reasoning rows accumulating, which is hard for a plain vector store to claim.
  • Publish everywhere at once. The Connectors section makes the same endpoint installable on frontier platforms, so distribution does not depend on your own site.

Connectors on frontier platforms

DGUI-HyperMem is a plain Streamable HTTP MCP server, so it can be attached to any platform that speaks remote MCP, and to any platform with OpenAI-compatible tool calling through a thin bridge. This section covers the two distribution paths that matter: the official MCP registry, which makes the server installable by name, and per-platform remote MCP configuration.

Publish to the MCP registry

The MCP registry is the closest thing MCP has to an app store: clients can list and install servers from it. The registry hosts metadata only, never artifacts, so the package has to exist somewhere public first — npm is the usual choice. Publishing is a four-step flow with the official mcp-publisher CLI.

1. Name the server in the package

Add an mcpName field to package.json. With GitHub authentication it must start with io.github.<your-account>/. The registry verifies that the published package and the registry metadata agree, so this value is what ties the two together.

2. Publish the package

Run npm publish --access public. The registry will not accept metadata for a package it cannot resolve.

3. Generate and edit server.json

Run mcp-publisher init to create the metadata template, then declare the remote endpoint. The name in server.json must match mcpName exactly. Validate without publishing using mcp-publisher validate.

4. Authenticate and publish

Run mcp-publisher login github, then mcp-publisher publish. Re-run the login command if a publish is rejected as an expired registry token.

Registry metadata for this server
{
  "$schema": "https://static.modelcontextprotocol.io/schemas/2025-09-29/server.schema.json",
  "name": "io.github.ctaxnagomi/dgui-hypermem",
  "description": "Hybrid long-term memory with a JEV reasoning layer, exposed over Streamable HTTP MCP.",
  "status": "active",
  "repository": {
    "url": "https://github.com/ctaxnagomi/dgui-hypermem",
    "source": "github"
  },
  "version": "1.0.0",
  "packages": [
    {
      "registryType": "npm",
      "identifier": "dgui-hypermem",
      "version": "1.0.0"
    }
  ],
  "remotes": [
    {
      "type": "streamable-http",
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp"
    }
  ]
}
Authentication is still per-user. Registry metadata advertises the endpoint, not a credential. Every caller supplies its own DGUI-HyperMem bearer token, so keep the hosted endpoint rate-limited and never publish an operator token in server.json.

Per-platform attachment

PlatformMechanismWhere it is configured
Codex CLI, IDE extension, ChatGPT desktopRemote Streamable HTTP MCP with a bearer token~/.codex/config.toml — see the Codex block
Grok via the xAI Responses APIRemote MCP tool in the request tools arrayRequest body — see the Grok block
ChatGPT webRemote MCP-backed tools supplied by a pluginPlugin manifest in your own project
Claude Code and Claude DesktopRemote MCP with bearer or OAuthCLI add command, or the desktop connector UI
VS Code, Cursor, JetBrains, Windsurf, ZedRemote MCP JSON with an Authorization headerWorkspace or user MCP settings
Any OpenAI-compatible endpointFunction calling against a bridge that wraps the MCP toolsYour application — see Integration JSON

Two constraints apply everywhere. Remote MCP must be reachable over the public internet with a valid TLS certificate, which the workers.dev endpoint satisfies and a raw *.workers.dev preview will not. And Codex reads the MCP instructions field during initialization and uses it as server-wide guidance, so this server's help text is what teaches a client when to reach for memory — the first 512 characters of that text matter most for tool selection.

Self-host DGUI-HyperMem

Deploy your own Worker so memory data, D1 records, Vectorize vectors, AI provider usage, and optional Hugging Face exports remain under your Cloudflare account.

1. Clone and install

Repository setup
git clone https://github.com/ctaxnagomi/dgui-hypermem.git
cd dgui-hypermem
npm install
npx wrangler login

2. Create Cloudflare resources

D1 and Vectorize
npx wrangler@latest d1 create dgui-hypermem
npx wrangler@latest vectorize create dgui-hypermem --dimensions=768 --metric=cosine

Copy the generated D1 database ID and Vectorize index name into wrangler.jsonc. Keep the binding names DB and VECTORIZE, or update src/types.ts and all binding references consistently.

3. Apply the schema

D1 migrations
npx wrangler@latest d1 migrations apply dgui-hypermem --remote
Current repository caveat: the checked-in migration chain predates several columns and tables queried by the current Worker, including token plan/training/connection fields, CRM device logging, and admin_logs. Before a first production deploy, reconcile a forward migration with the schema expected by the current source. Do not treat an unmodified fresh migration run as production-ready.

4. Configure secrets and bindings

  • Set MCP_TOKEN, ADMIN_PASSKEY, and any TypeSafe AI or Hugging Face secrets with Wrangler rather than committing them.
  • Bind D1 as DB, Vectorize as VECTORIZE, and Workers AI as AI.
  • Use the 768-dimension @cf/baai/bge-base-en-v1.5 model with a 768-dimension Vectorize index.
  • Set DEFAULT_SCOPE and optional JEV mode variables to match your deployment policy.

5. Validate and deploy

Typecheck, deploy, and verify
npm run typecheck
npm run deploy
curl "https://YOUR_WORKER_HOST/health"

After deployment, create an initial CRM token through your approved operator workflow, call /api/verify-token, and complete an add/search/forget smoke test before connecting clients.

Security and data handling

  • Set MCP_TOKEN; the current Worker permits unauthenticated access when it is absent.
  • Set ADMIN_PASSKEY as a secret and remove passkey values from source, Wrangler variables, examples, and documentation.
  • Prefer wrangler secret put over plaintext vars for tokens, passkeys, Stripe secrets, and Hugging Face credentials.
  • Do not log authorization headers, API keys, access tokens, private dataset tokens, or passkeys.
  • Treat memory content as retained data. Apply retention, deletion, and tenant-isolation requirements at the application layer.
  • Disable or tightly control JEV dataset export when memory content must not leave the primary deployment.
  • Use HTTPS-only Worker routes and restrict administrative endpoints at the network or application layer.
Current deployment audit: the repository has known hardening gaps around fail-open authentication, source-level secret fallback, setup-token logging, and migration completeness. Review and remediate them before using a fresh self-hosted deployment for sensitive or multi-tenant data.

Troubleshooting

SymptomLikely causeResolution
401 unauthorizedMissing, malformed, inactive, or incorrectly scoped tokenSend Authorization: Bearer TOKEN and call /api/verify-token.
429 quota exceededToken request allowance is exhaustedInspect /api/check-quota and wait for resets_at or ask an operator to adjust the plan.
MCP client cannot connectWrong URL, missing header, or client-specific schema mismatchUse the exact client block below and confirm the URL ends in /mcp.
Search returns no resultsDifferent scope, wrong type filter, or durable_only filteringRepeat with the same scope and no restrictive filter, then call list.
Vector search is unavailableAI binding, Vectorize binding, dimensions, or model configuration is wrongConfirm AI and VECTORIZE and use a 768-dimension index.
D1 reports a missing column or tableFresh migration schema is behind the current WorkerApply a reconciled forward migration before deploying the current source.
JEV sync returns zero or failsProvider is off, dataset secret is absent, or queue rows are emptyCall jev_queue_stats, inspect Worker logs, and verify Hugging Face configuration.

Agentic layer

Registering the server is not the same as an agent being able to use it. A client can hold a valid configuration and still end up with an inert model: tools filtered out of the catalog, approvals blocking every call, or a tool result truncated before the model reads it. The SDK configuration layer exists to close that gap, so the same endpoint behaves like a working agentic toolset on any frontier platform.

Universal agentic execution layer

Whichever platform or SDK sits in front of the model, the layer is responsible for six things. Treat this as the contract to verify before declaring an integration live.

GuaranteeWhat it meansHow to verify
Full tool catalogAll eight tools reach the model, with no allowlist silently dropping one.List the tools the client reports after connecting and confirm the count.
Tool calling enabledThe model is permitted to emit tool calls, with the choice left to the model rather than pinned to a single function.Send a request that can only be answered by a tool and confirm a call is emitted.
No blocking approvalApprovals do not stall an unattended run.Run headless and confirm a tool call completes without a prompt.
Untruncated resultsTool output is not clipped by a per-tool token budget smaller than a real result.Store a long memory and confirm the recalled text arrives complete.
Server instructions surfacedThe MCP instructions field reaches the model as system-level guidance.Ask the model what the server is for without quoting the manual.
Multi-turn tool loopResults feed back and the model can call again, so recall-then-write works in one turn.Search, act on the result, then add, in a single conversation.

The Worker side already satisfies its half of this contract. It returns server instructions during initialization, exposes all eight tools, applies no output cap of its own, and allows a search limit of up to 50 results. Everything below is therefore about client configuration.

SDK configuration per platform

PlatformTool-calling switchOutput-capability switch
Codexrequired = true, enabled_tools, default_tools_approval_mode = "auto"tools.<tool>.output_token_limit, plus tool_timeout_sec
Grok via xAIallowed_tools in the MCP tool entry; omit it to inject every toolNo per-tool cap in the MCP entry; bound the loop with max_tool_iterations
Claude CodeTools are auto-discovered from the connected serverClient default; no per-tool key
VS CodeDeclared in servers; confirm no allowlist is filtering the catalogClient default; no per-tool key
Cursor, JetBrains, Hermes, OpenClawDeclared in the MCP JSON; confirm no allowlist is filtering the catalogClient default; no per-tool key
opencodeenabled: true on the server entryClient default; no per-tool key
Custom SDKForward the full tool schema array and leave tool_choice unset or "auto"Raise or remove your own per-tool output cap before the model sees the result
Agentic layer brief
You have DGUI-HyperMem tools available through MCP. Treat them as part of your
own capability set, not as an optional add-on.

Tool calling
- The tools are add, search, list, profile, forget, help, sync_jev_dataset, jev_queue_stats.
- Call them yourself. Do not describe the tools to the user and stop.
- Prefer searching memory before answering, and storing a memory after a durable decision.

Output
- Use the full result you are given. Do not summarise away ids, tags, or scores you may need.
- If a result is empty, say so and continue rather than retrying the same call in a loop.

Boundaries
- Never store credentials, private keys, payment data, or unnecessary personal information.
- Scopes organise memory; they are not authorization boundaries.

Copy for LLM

Give an agent the Markdown brief and exactly one client block. Replace placeholders through environment variables or secret prompts. Never paste a live token into a conversation or tracked file.

Markdown brief

Markdown
# DGUI-HyperMem integration

Use the DGUI-HyperMem remote MCP server for durable memory.

- Endpoint: https://dgui-hypermem.ctaxnagomi.workers.dev/mcp
- Transport: MCP Streamable HTTP
- Authentication: Authorization: Bearer DGUI_HYPERMEM_TOKEN
- Default hosted scope: default
- Tools: add, search, list, profile, forget, help, sync_jev_dataset, jev_queue_stats

Before starting work:
1. Search the relevant project or user scope.
2. Use returned memories as context, but verify them against the current repository and user request.
3. Do not store credentials, private keys, payment data, or unnecessary personal information.

After a durable decision:
1. Add one concise memory with a stable scope, tags, and source.
2. Include constraints and rationale when they matter.
3. Correct stale records by forgetting the old ID and adding the replacement.

When deleting:
- Prefer forget with exact IDs.
- Use all:true only after confirming the intended scope.

Scopes organize memories but are not tenant authorization boundaries in the current hosted implementation.

Generic MCP JSON

Generic mcpServers
{
  "mcpServers": {
    "dgui-hypermem": {
      "type": "http",
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "headers": {
        "Authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}"
      }
    }
  }
}

Portable integration JSON

Integration manifest
{
  "name": "dgui-hypermem",
  "version": "1.0.0",
  "transport": "streamable-http",
  "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
  "authentication": {
    "type": "bearer",
    "header": "Authorization",
    "environmentVariable": "DGUI_HYPERMEM_TOKEN"
  },
  "defaultScope": "default",
  "tools": [
    "add",
    "search",
    "list",
    "profile",
    "forget",
    "help",
    "sync_jev_dataset",
    "jev_queue_stats"
  ]
}

VS Code

Add this to .vscode/mcp.json. VS Code prompts for the input value instead of storing it in the file.

VS Code mcp.json
{
  "servers": {
    "dgui-hypermem": {
      "type": "http",
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "headers": {
        "Authorization": "Bearer ${input:dgui_hypermem_token}"
      }
    }
  },
  "inputs": [
    {
      "type": "promptString",
      "id": "dgui_hypermem_token",
      "description": "DGUI-HyperMem bearer token",
      "password": true
    }
  ]
}

Claude Code

Claude Code CLI
claude mcp add --transport http dgui-hypermem \
  https://dgui-hypermem.ctaxnagomi.workers.dev/mcp \
  --header "Authorization: Bearer ${DGUI_HYPERMEM_TOKEN}"

Cursor

Cursor mcp.json
{
  "mcpServers": {
    "dgui-hypermem": {
      "type": "http",
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "headers": {
        "Authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}"
      }
    }
  }
}

JetBrains

Use a project MCP JSON file such as .junie/mcp/mcp.json, or add the same server through your IDE’s MCP configuration UI.

JetBrains MCP JSON
{
  "mcpServers": {
    "dgui-hypermem": {
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "headers": {
        "Authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}"
      }
    }
  }
}

Hermes

Hermes mcp_servers
{
  "mcp_servers": {
    "dgui-hypermem": {
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "headers": {
        "Authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}"
      }
    }
  }
}

OpenClaw

OpenClaw MCP config
{
  "mcpServers": {
    "dgui-hypermem": {
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "transport": "streamable-http",
      "headers": {
        "Authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}"
      }
    }
  }
}

opencode

opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "dgui-hypermem": {
      "type": "remote",
      "url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "enabled": true,
      "oauth": true
    }
  }
}

Codex

Codex reads config.toml, shared by Codex CLI, the IDE extension, and the ChatGPT desktop app. The settings below are the agentic layer for Codex: required = true makes startup fail loudly instead of silently dropping the server, enabled_tools pins the full catalog, default_tools_approval_mode = "auto" stops an unattended run from stalling on a prompt, and output_token_limit overrides the model’s default per-tool truncation so a large recall is not clipped.

Codex config.toml
# ~/.codex/config.toml

[mcp_servers.dgui-hypermem]
url = "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp"
bearer_token_env_var = "DGUI_HYPERMEM_TOKEN"
required = true
startup_timeout_sec = 20
tool_timeout_sec = 60
default_tools_approval_mode = "auto"
enabled_tools = ["add", "search", "list", "profile", "forget", "help", "sync_jev_dataset", "jev_queue_stats"]

[mcp_servers.dgui-hypermem.tools.add]
approval_mode = "approve"
output_token_limit = 4000

[mcp_servers.dgui-hypermem.tools.search]
approval_mode = "approve"
output_token_limit = 8000

[mcp_servers.dgui-hypermem.tools.list]
approval_mode = "approve"
output_token_limit = 8000

[mcp_servers.dgui-hypermem.tools.profile]
approval_mode = "approve"
output_token_limit = 4000

[mcp_servers.dgui-hypermem.tools.forget]
approval_mode = "approve"
output_token_limit = 2000

[mcp_servers.dgui-hypermem.tools.help]
approval_mode = "approve"
output_token_limit = 2000

[mcp_servers.dgui-hypermem.tools.sync_jev_dataset]
approval_mode = "approve"
output_token_limit = 2000

[mcp_servers.dgui-hypermem.tools.jev_queue_stats]
approval_mode = "approve"
output_token_limit = 4000
Two supporting switches. Set mcp_optional_startup_grace_ms = 0 at the top level if a cold start is losing tools from the initial catalog, and use codex mcp list to confirm the server and all eight tools are present. Codex also reads the MCP instructions field as server-wide guidance, so keep the first 512 characters of this server’s help text self-contained.

Grok

Grok reaches the server through the xAI remote MCP tool, declared in the tools array of the request. Listing every tool in allowed_tools is what keeps the whole catalog in front of the model; omitting the field also injects everything, but stating it explicitly keeps the agentic layer auditable. tool_choice is left on auto so the model decides when memory is relevant instead of being pinned to one function.

xAI Responses API request
{
  "model": "grok-4.7",
  "tool_choice": "auto",
  "tools": [
    {
      "type": "mcp",
      "server_url": "https://dgui-hypermem.ctaxnagomi.workers.dev/mcp",
      "server_label": "dgui-hypermem",
      "server_description": "Hybrid long-term memory with a JEV reasoning layer. Use search before answering and add after a durable decision.",
      "authorization": "Bearer ${DGUI_HYPERMEM_TOKEN}",
      "allowed_tools": [
        "add",
        "search",
        "list",
        "profile",
        "forget",
        "help",
        "sync_jev_dataset",
        "jev_queue_stats"
      ]
    }
  ]
}
SDK naming differs from the REST body. In the xAI native SDK the same tool is mcp(server_url=..., server_label=...), and the two fields are renamed: allowed_tools becomes allowed_tool_names, and headers becomes extra_headers. Only Streamable HTTP and SSE transports are supported, and the OpenAI-specific require_approval and connector_id parameters are not.

ctecx-instruct task log

Paste this into an agent to make it log work as a ctecx_instruct pack, the five-part format defined in ctaxnagomi/ctecx-instruct.

ctecx_instruct brief
CTECX task logging is enabled for this work.

Repository: https://github.com/ctaxnagomi/ctecx-instruct
Format: ctecx_instruct@1, a five-part, zip-able instruction pack.

For every task, produce one pack with these five files at the zip root, named
ctecx_instruct_<task_id>.zip, where task_id is ctecx-<topic>-<seq>.

INSTRUCT.md
  - Task log with the required blocks: Objective, Important Details, Work State
    (Completed / Active / Blocked), and Next Move, plus Execution Steps,
    Deliverables, and Verification.
  - Keep it resumable: a fresh session must be able to continue from the file alone.

task.sh
  - Start with set -euo pipefail.
  - Idempotent and safe by default; gate destructive commands behind RUN_DESTRUCTIVE=1.
  - Expose setup / build / run / test stages; test must exit non-zero on failure.

task.sql
  - SQLite schema that creates agent_memory (msg_id, task_id, sender, role, payload,
    embedding, ts) and task_meta (task_id, owner, status, started_at, closed_at),
    plus a task_audit trail.

task.json
  - format_version "ctecx_instruct@1", task_id, owner, created, status, params, tags, parts.
  - Ship a manifest with sha256 for INSTRUCT.md, task.sh, task.sql, task.assembly.
    task.json is excluded from its own manifest.

task.assembly
  - The plan as a program listing: labelled segments that mirror the INSTRUCT steps,
    register-style state, and opcodes MOV, CALL, CMP, JZ, RET, HLT.

Rules
  - Never embed secrets, credentials, or private keys in any part.
  - Every constraint the executor must not violate belongs in Important Details.
  - Do not close a task until Verification passes.
  - Keep the DGUI-HyperMem integration contract: use the remote MCP server for
    durable memory and never store credentials in memory.