- TypeScript 95.9%
- Shell 3.1%
- PowerShell 0.6%
- JavaScript 0.4%
Six rough edges reported from real use of the base tools.
Every problem in a row is reported. toCells converted values until the
first failure, so a row with a bad "Operation type" and a bad
"Hypervisor" reported only the first and took two retries. It now
converts every value and throws one CellProblems carrying the list; bulk
results carry `errors: string[]` per row.
list_base_rows returns real types. Values were CSV export's strings:
"3", "true", a multi-select as "26, 25" — ambiguous as soon as an option
label contains a comma. typedCell returns numbers, booleans, labels,
arrays, { id, name } / { id, title } for people and pages, ISO dates,
and null for empty. Every shape is accepted back by the writes.
An unset checkbox reads false, not "".
update_property no longer contradicts itself. The tool said renaming an
option "updates every row", the field said "does not touch any row".
Both true — rows store option ids — and now worded as one fact.
Views echo their real setup. They showed only `hasFilter: true` and
never visibleProperties. They now report the effective `columns` in
display order, `hiddenProperties`, `sorts`, and the `filter` translated
back into the where/match vocabulary, so it can be passed straight back;
a UI-built nested group stays nested. get_base includes this for every
view. Columns are resolved exactly as the client resolves them.
visibleProperties is now written as the table UI writes it: the hidden
complement plus a column order, with visiblePropertyIds nulled. The
client always saves that form and lets the hidden list win, so a
whitelist was silently reinterpreted the first time anyone touched the
view. The filter vocabulary gains on_or_before, on_or_after, has_all and
is_within, the four operators the UI can build that it lacked.
trash_page moves a page and its children to the space trash, mirroring
core's POST /pages/delete without permanentlyDelete, including the
PAGE_TRASHED audit event. On a base it is the restorable way to remove
it. It refuses a base row's own page. No permanent delete is offered.
Also fixed: get_page on a row page resolved people and pages using the
references of the base's FIRST row (borrowed from list({limit: 1})), so
names on every other row came back empty. It now resolves its own row;
BaseRowService.buildReferences is public for that.
read-shapes.spec.ts adds 22 tests, base-schema-tools.spec.ts 2. Negative
controls (blank checkbox, first error only, whitelist views, joined
multi-selects) each fail the suite when reintroduced.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|---|---|---|
| .github/workflows | ||
| api-key | ||
| attachments-ee | ||
| audit | ||
| base | ||
| docker | ||
| document-import | ||
| docx-export | ||
| group-role | ||
| licence | ||
| lockfile | ||
| maintenance | ||
| mcp | ||
| mfa | ||
| oauth | ||
| openapi | ||
| page-move | ||
| page-permission | ||
| page-verification | ||
| patches | ||
| pdf-export | ||
| personal-space | ||
| scim | ||
| scripts | ||
| shared | ||
| sso | ||
| template | ||
| typesense | ||
| upstream-contract | ||
| .env.example | ||
| .gitattributes | ||
| .gitignore | ||
| compose.dev.yml | ||
| compose.example.yml | ||
| DEPLOYMENT.md | ||
| di-check.ts | ||
| docker-compose.ee.yml | ||
| ee.module.ts | ||
| INSTALL.md | ||
| ISSUES.md | ||
| README.md | ||
| SETUP-DEV.md | ||
| TODO.md | ||
| upstream.pin | ||
| VERIFICATION.md | ||
Docmost Freenterprise — EE bundle
A self-hosted replacement for Docmost's private ee submodule
(github.com/docmost/ee), which is closed-source and unavailable.
This directory is its own git repository
(ssh://git@git.derg.cz/ulysia/docmost-freenterprise.git), checked out into
the parent Docmost fork at apps/server/src/ee/ and gitignored there.
Why this works at all
Docmost's OSS core is built to load enterprise code that it doesn't ship. Two mechanisms:
1. Scattered require() hooks. Core calls into EE at specific points,
each wrapped in try/catch:
// apps/server/src/core/auth/strategies/jwt.strategy.ts
ApiKeyModule = require('./../../../ee/api-key/api-key.service');
const ApiKeyService = this.moduleRef.get(ApiKeyModule.ApiKeyService, { strict: false });
return ApiKeyService.validateApiKey(payload);
Each hook has a documented fallback when the module is absent — silent no-op,
thrown error, or a "requires a valid enterprise license" message. Every
module in this bundle documents its own hook contract in its header comment.
Read those first; they cite exact file:line call sites in core.
2. The root aggregator. app.module.ts does
require('./ee/ee.module')?.EeModule and pushes it into the root imports.
Anything registered under EeModule becomes resolvable to every
ModuleRef.get(..., { strict: false }) hook above, and any controller it
mounts gets routed.
The key insight
Core already ships the schema, and usually the client, for every EE
feature. All 39 tables have migrations; there is no gap between the schema
Kysely knows about and what migrations create. So almost nothing here needed a
new migration — user_mfa, api_keys, audit, auth_providers,
scim_tokens, page_verifications, base_* all already existed.
The client is frequently complete too. When implementing a feature, the API
contract is usually already sitting in apps/client/src/ee/<feature>/ —
reverse-engineer it from there rather than inventing routes. Several
non-obvious route names (/docx-export, /search-attachments,
/pages/verification-info) are fixed by client code we don't control.
Status
| Feature | State | Notes |
|---|---|---|
| SSO — OIDC | implemented | Authentik-targeted. Local-account merge via SSO_ACCOUNT_MERGE; group-claim sync via SSO_OIDC_GROUP_SYNC. SAML/Google/LDAP are stubs. |
| MFA | implemented | TOTP + backup codes, via otpauth. |
| API keys | implemented | JWT-based; no separate secret stored. |
| Audit logging | implemented | Real IAuditService over the audit table. |
| Licence | implemented | Always licensed; edition shown as "Freenterprise". |
| PDF export | implemented | Gotenberg renders the real page in Chromium. |
| PDF import | implemented | Tika text extraction. |
| DOCX export | implemented | Reuses the OSS prosemirror-docx serializer. |
| DOCX import | implemented | mammoth; images become real attachments. |
| Attachment search | implemented | Tika indexing + /search-attachments. |
| SCIM 2.0 | implemented | Users + Groups, deprovisioning, group sync. |
| OpenAPI docs | implemented | Swagger UI at /openapi for all 223 routes — see openapi/. |
| MCP server | implemented | /mcp over streamable HTTP, tools on core services — see mcp/. |
| OAuth 2.1 provider | implemented | PKCE, dynamic registration, refresh rotation — see oauth/. Backs MCP auth. |
| Group-driven roles | implemented | A group can grant member/admin/owner to its members — see group-role/. |
| Page verification | implemented | Both expiring and qms workflows. |
| Bases / Kanban | implemented | All six stages — see base/base.module.ts. |
| Typesense | stub | Deliberately deferred — see below. |
| Confluence import | dropped | Deleted; not needed. |
Bases, in six stages
All implemented, in this order — each left the feature more usable than the last, and the risky parts came after the basics worked:
- CRUD — 23 endpoints. A base is a page with
is_base = true, so it inherits the page tree, space membership and page-level permissions. - Realtime —
BaseWsService, 14 outboundbase:*events. Broadcasts include the originating client, which suppresses its own echo byrequestId. - Query engine (
base/engine/) — compiles the client'sFilterNodeinto SQL over thebase_cell_*helpers. 20 operators, timezone-aware relative date presets. - Formulas — values are materialised into
cells, which is what lets stage 3 filter and sort them. Source is recompiled server-side; the client's AST and dependency list are discarded. - Async type conversion — staged via
pending_type/pending_token, applied by a worker, committed with abase_schema_versionbump. - CSV export.
BaseProcessor is the consumer for BASE_QUEUE, which core registers but
never drains.
Deliberately not done
- Typesense — Postgres FTS here is already competent (weighted tsvector,
unaccent,pg_trgm, GIN). Typesense would add a second stateful service plus an indexing pipeline and query-time ACL filtering, for typo tolerance and scale we don't currently need. LeavingSEARCH_DRIVERunset is the safe default; setting it totypesensewithout the module hard-fails search. - AI search (semantic / vector) — deferred for the same reasons as Typesense, more strongly. See below.
- AI chat — needs the same embedding pipeline as AI search, plus a
streaming tool-call loop.
ai_chats/ai_chat_messagesalready exist in core, and the whole client ships inapps/client/src/ee/ai-chat/. - SAML / Google / LDAP — stubs under
sso/strategies/. Schema and client forms exist if ever needed.
Typesense and vector search are the same project twice
Worth writing down because it is not obvious from the outside: these two look like different features and are mostly the same work. Whichever gets built first pays for most of the second.
Both require a second index that Postgres FTS does not maintain, and therefore all of:
- an indexing pipeline hooked to page create / update / delete / move, run through a queue rather than inline;
- a backfill for existing content, and a reconciliation path for when the index drifts;
- query-time ACL filtering — neither index knows Docmost's permissions, so results must be constrained to the caller's accessible spaces and pages. This is the security-critical part and the reason neither is a weekend job: get it wrong and search leaks page titles or excerpts across space boundaries. A page moving space changes who may see it without its content changing at all.
Where they differ, and it does not favour Typesense as much as it first looks:
| Typesense | Vector (pgvector) | |
|---|---|---|
| Extra service | yes, stateful | no — same Postgres |
| Unit indexed | page | chunk (so more rows, and a chunking strategy) |
| Per-edit cost | one HTTP call | an embedding provider call per chunk |
| Failure mode | stale results | stale results and a bill |
| Model changes | none | changing the embedding model means re-embedding everything |
So AI_VECTOR_DRIVER=pgvector is actually cheaper than Typesense on the
"another daemon to run" axis, and more expensive on the write path. Core
already ships the whole configuration layer for it — AI_DRIVER,
AI_VECTOR_DRIVER, AI_EMBEDDING_MODEL, AI_EMBEDDING_DIMENSION,
OPENAI_API_KEY, and the rest, with validation — but not the
page_embeddings table it probes for in isPageEmbeddingsTableExists(). That
table is the EE side's to create, which here means a new patch.
Generative AI (/ai/generate) is the exception to all of the above: it needs
no index, no pipeline and no embeddings — just a provider call and a stream —
so it is separable from both, and separately gated by the generativeAi
setting.
SSO local-account merge
SSO_ACCOUNT_MERGE controls whether an SSO login adopts an existing local
account instead of creating a duplicate:
| Value | Behaviour |
|---|---|
email |
match on email (default — what the bundle always did) |
email-or-username |
also match preferred_username against users.name |
off |
never adopt; a colliding email is a clear error |
Env var rather than a UI toggle because the client's IAuthProvider has no
such field and adding one would mean editing client code we don't control.
Guards, since this is an account-takeover surface:
- Email merge requires
email_verified: true. A missing claim counts as unverified — treating absence as good enough would trust any IdP that simply omits it. Authentik sends it. SetSSO_MERGE_REQUIRE_VERIFIED_EMAIL=falseto relax this to "not explicitly false" for a trusted IdP that genuinely doesn't. - Username merge requires exactly one local match, since
users.namecarries no unique constraint. Ambiguous matches are logged and skipped. - Disabled accounts are never adopted, and every merge is audit-logged with which field matched.
Note email-or-username is inherently the weaker mode: it links on a claim
the user may control at the IdP, matched against a non-unique display name.
Prefer plain email unless you specifically need it.
Conventions
- Never edit core files. Every line added to the parent repo is future
rebase conflict surface. New env vars go in
shared/ee-env.tsreadingprocess.env, not into core'sEnvironmentService. (Vars core already defines —GOTENBERG_URL,APP_URL— still go throughEnvironmentService.) - Match core's conventions rather than inventing parallel ones: same
tsquery/ts_rank/f_unaccentsearch construction, same cursor pagination shape, same CASL ability checks, samecatch (err: any)style. - Reuse core services —
SignupService,PageAccessService,SessionService,TokenService, the repos. Don't reimplement permissions. - Most core modules are
@Global()(Database, Environment, Storage, Queue, Casl, PageAccess), so sub-modules usually need to import onlyTokenModuleorAuthModule.
Documentation
| INSTALL.md | Deploying with Docker, start to finish. The documented path. |
| SETUP-DEV.md | Running from source, for working on the code. |
| DEPLOYMENT.md | The reference behind both — upstream pin, the two-image story, storage, CI. |
| .github/workflows/build.yml | Runs check-upstream.sh on main, testing and every PR; publishes the image only if it passed, and only from main. |
| patches/README.md | What each client patch fixes, and how to regenerate one. |
| upstream-contract/README.md | Unit tests that fail when upstream moves a hook or a SQL function. Run after changing the pin. |
| VERIFICATION.md | What has and hasn't actually been tested. |
| ISSUES.md | Bugs found in live testing. |
| TODO.md | Wanted, not built. |
Templates live here too: compose.example.yml (pulls the published image), compose.dev.yml (builds from source) and .env.example.
Installing
CI publishes the image, so a deployment needs neither the source nor a toolchain — two files and a directory:
.
|-- compose.yml <- from compose.example.yml
|-- .env <- from .env.example
`-- data/ <- attachments, owned by uid 1000
INSTALL.md opens with a single paste-able block: edit
APP_URL, run it, and you have both files with generated secrets and a
correctly owned storage directory. Then docker compose up -d.
The image is git.derg.cz/ulysia/docmost-freenterprise:main, rebuilt on every
push here. main moves; every build also publishes a :<short-sha> tag to pin
to, and labels each image with the upstream Docmost commit it was built
against — the version that matters is in neither repository alone.
Building it yourself instead needs both clones and the patch machinery; that is
the Building from source section of INSTALL.md, with compose.dev.yml.
One thing worth knowing before the first start: data/ must exist and be owned
by uid 1000. The container runs as node, and Docker creates a missing
bind-mount source as root — which fails every upload with a message that says
nothing useful.
mcp
/mcp — a Model Context Protocol server, so an AI client can work with the
workspace directly instead of through an integration layer.
Enable it in Settings → AI → MCP (workspace.settings.ai.mcp). Off by
default; the endpoint 404s until it is on, so a workspace that has not opted in
is not merely unauthenticated but absent.
Connect by pointing an MCP client at https://<host>/mcp. Nothing is
configured by hand: the client gets a 401 carrying
WWW-Authenticate: Bearer resource_metadata="…", follows it to the discovery
documents, registers itself, and sends you to the consent screen. See oauth/.
Auth is core's JwtAuthGuard, so either an OAuth access token or an API
key works — upstream's semantics, which its own settings panel states. Turning
on Enforce OAuth (settings.ai.enforceMcpOauth) rejects API keys.
Tools
Thirty-one. The page, space and comment tools mirror the core routes upstream
annotated with @OAuthScope — that annotation set is upstream's statement of
what an OAuth token is meant to reach, so mirroring it keeps the boundary
aligned rather than invented. The base tools go further, because bases are
ours and upstream has no annotations for them to mirror.
| Tool | Scope | Notes |
|---|---|---|
whoami |
read | Identity and granted scopes |
list_spaces |
read | Only spaces the user belongs to |
search_pages |
read | Full-text, permission-filtered by core |
get_page |
read | Content converted to markdown |
list_recent_pages |
read | |
list_page_children |
read | Walks the page tree |
create_page |
write | Markdown in; builds the Yjs doc correctly |
update_page_content |
write | append / prepend / replace |
rename_page |
write | Title and icon. Renames a base too — its name is its page title |
list_bases |
read | Bases in a space |
get_base |
read | Columns, types, options with ids, and each view's full setup |
list_base_rows |
read | Values keyed by property name, as real types; simple filters |
get_space |
read | |
list_comments |
read | Flattened to text |
list_page_attachments |
read | |
create_comment |
write | Markdown; replies via parentCommentId |
move_page |
write | before/after a sibling, first/last, or another space. Moves a base with its rows; refuses a base row's own page |
trash_page |
write | To the space trash, restorable; works on a base. No permanent delete |
create_base_row |
write | Values by property name, options by label |
update_base_row |
write | Partial patch; null clears a value |
delete_base_row |
write | |
create_base_rows |
write | Bulk; every problem per row, the rest still land |
upsert_base_rows |
write | Create or update, matched on a key column |
create_base |
write | Base plus its columns in one call |
delete_base |
write | Destroys rows, columns and views; the page survives. trash_page removes a base restorably |
create_property |
write | Options, includeTime, numberFormat |
update_property |
write | Rename, settings, add/rename/remove/reorder options |
delete_property |
write | Not the primary column |
create_base_view |
write | Sort, filter, visible/hidden columns, kanban grouping |
update_base_view |
write | Patch; empty array clears sorting or filtering |
get_page also reports isBase, and for a page that is a base row it
returns that row's values alongside the content. Without that a row page
reads as empty, because its field values live in base_rows.cells and were
never part of the page — which is what made bases look unreachable over MCP.
Bases over MCP
Cells are stored keyed by property id, with values that are ids too
(opt_3f2a for a select). Unusable by a model in both directions, so the
tools translate at the boundary:
- reading — values come back keyed by property name, as their real
type (
typedCellinmcp/services/base-tool-shared.ts): numbers, booleans, option labels, arrays for multi-selects,{ id, name }for people,{ id, title }for pages, ISO 8601 UTC for dates,nullwhen empty. An unset checkbox isfalse. These used to be strings from CSV export'srenderCell—"3","true", a multi-select as"26, 25"— which made a sync's comparisons unreliable and a comma inside an option label ambiguous. Every shape a read returns is accepted by the writes, so a row can be read, edited and written back without reformatting. - writing — callers pass names and labels, and
toCellValuemaps them back to ids. An unmatched option fails with the valid list rather than writing the label as a plain string, which would render as itself and silently not be the option anyone meant.
Person and page cells take ids only. Resolving a name would mean guessing between duplicates, and guessing which person a value refers to is not a thing to do quietly.
Permissions come free: BaseService.getInfo runs through PageAccessService
because a base is a page, BaseRowService.list does its own
requireViewableBase, and its reference resolver is permission-filtered — a
cell pointing at a page the caller cannot see comes back empty rather than
leaking a title.
Write tools are not registered at all on a read-only grant, so a model plans against what it can actually do rather than discovering the refusal mid-task.
Changing a base's shape
The schema tools live in mcp-base-schema-tools.service.ts, separate from the
row tools: one translates cell values, the other translates a schema, and
neither belongs in the other's file. Both resolve their vocabulary through
base-tool-shared.ts, so a read and a write cannot disagree about what a
property looks like.
Names or ids, everywhere. Every argument that takes a property, an option or a sort column accepts either. Names are what a prompt can state; ids survive a rename and are what a sync should persist. Ids are tried first — a property literally named like another one's id is rare, but the order has to be fixed rather than incidental. Two properties sharing a name is an error naming both ids, never a guess.
Every property change returns the whole schema, option ids included. A
generated option id is otherwise unknowable until the next get_base.
Renaming an option touches no rows. Removing one clears it from every row.
Cells store option ids, never labels, so a rename is a one-row update to the
property and every cell stays correct by construction — fixing "Maintanence"
in place is instant and safe. A removal is the opposite: BasePropertyService
ripples it into every cell that held it (engine/prune-choices.ts), and
update_property reports how many with clearedCells, counted with the same
pruneCellValue the service then runs, so the number is what actually
happens rather than an estimate. Edits apply in a fixed order — rename,
remove, add, reorder — so order can name an option the same call renamed or
added, and remove-then-add of one label is a replacement rather than a
duplicate.
typeOptions is replaced, not merged, on the server. Every option edit is
therefore read-modify-write against the current property; a partial object
would silently discard every setting it omitted.
Dates are UTC instants, and a date-only column is pinned to noon. The
client renders a stored instant with local getters, so midnight UTC shows as
the previous day anywhere west of Greenwich. With includeTime: false, a
written date keeps its UTC calendar day and is stored at 12:00Z — the one
hour that survives every real timezone. With includeTime: true the instant is
stored exactly, offsets converted to UTC. get_base states which applies to
each date column (includeTime, storedAs). Before this, toCellValue passed
dates through raw, so "2026-03-01" was stored as that string and parsed by
the client as midnight UTC.
Bulk writes are per row, not per batch. create_base_rows and
upsert_base_rows report each row's outcome with its index in the request. One
bad option label fails that row; an all-or-nothing transaction would roll back
28 good rows over one typo and leave the caller to work out which. A failed row
lists every problem it has, not just the first: all its values are
converted before anything is reported, so one retry can fix the lot.
Views are written the way the table UI writes them. The client saves
hiddenPropertyIds plus propertyOrder and nulls visiblePropertyIds, and
when both are present the hidden list wins. A whitelist written over MCP would
render the same and then be reinterpreted the first time anyone touched the
view in the browser. So visibleProperties becomes the hidden complement, and
its order becomes the column order. Every view the tools return — and every
view in get_base — reports the effective setup, resolved as the client
resolves it: columns in display order, hiddenProperties, sorts, and the
filter translated back into the same where/match vocabulary the tools
take, so it can be passed straight back. A nested group built in the UI comes
back nested rather than flattened into a different filter.
Upsert matches exactly. After trimming, case-sensitively: VM-01 and
vm-01 are two machines until someone says otherwise. It reads the key column
once (paging the whole base) rather than querying per row, and a key repeated
within one batch updates the row that batch just created. A key that is
already duplicated in the base is a hard error listing the offending row ids —
there is no correct row to update, and picking one would corrupt the other on
every later sync. Multi-value columns (multiSelect, person, page, file) cannot
be keys.
What these tools deliberately do not do:
- Change a column's type. A conversion stages a background job that rewrites every cell and bumps the schema version; a tool returning before it finished would hand back a schema about to change underneath the caller.
- Change which column is primary. The UI does not expose it either, so it
would need a client patch;
create_basetakesprimaryNameto rename the seeded primary, which is a different thing. - Permanently delete anything.
trash_pagemirrors core'sPOST /pages/deletewithoutpermanentlyDelete: edit permission, thenPageService.removePage, then the samePAGE_TRASHEDaudit event. On a base it is the restorable way to remove it — rows, columns and views stay attached to the trashed page.delete_baseis the destructive one. Permanent deletion needs space-admin rights and cannot be undone, so it stays in the UI. - Move rows between bases. Not built; for a split, import fresh.
Two things that are easy to get wrong here
Permissions are core's, not ours. Every tool runs as the authorizing user
and calls the same PageAccessService / CASL checks the matching controller
does — including the subtle one, where creating a child page requires edit on
the parent but creating a root page requires Create on the space. A second
permission implementation would be a second thing to keep correct, and the
failure mode is handing an LLM someone else's private pages.
Content writes go through the collaboration gateway. Page bodies are Yjs
documents, so PageService.updatePageContent hands the change to
CollaborationGateway rather than writing the row. A direct write would be
silently overwritten by any connected editor.
Scope is enforced twice, deliberately: core's @OAuthScope('read') on the
route is the door policy (one HTTP request carries many operations, so the most
it can assert is the minimum), and each write tool re-checks for write.
oauth
An OAuth 2.1 authorization server, existing to authenticate MCP clients.
Core provides the hook — jwt.strategy requires
ee/oauth/services/oauth-strategy.service and calls
validateOAuthToken(payload, { workspaceId, host }) — plus the whole consent
UI (apps/client/src/ee/oauth/), the four oauth_* tables, and the route
exclusions in main.ts. The schema is what specifies the design; this module
implements it rather than choosing it.
| Scopes | read, write — fixed by upstream's consent screen |
| Grants | authorization_code (PKCE, S256 only), refresh_token |
| Access token | JWT, type: oauth_access, jti recorded in oauth_tokens |
| Refresh token | Opaque, stored as SHA-256, rotated on every use |
| Consent | One row per (user, client) in oauth_grants, revocable |
| Registration | RFC 7591, open — see below |
| Discovery | RFC 8414 + RFC 9728 at the three .well-known paths |
Revocation is immediate, not expiry-bound. A JWT cannot be un-issued, so every request re-checks the token row, the grant, and the user against the database. Without that, "revoke" would mean "revoke within the hour", which is not what the button says.
Replay fails safe. A reused authorization code, or a refresh token that has already been rotated away, revokes every token on that grant rather than just being refused — RFC 6819's advice. We cannot distinguish a buggy client from a stolen credential, and the honest client loses only one re-authorization.
Open registration is not a hole. Anyone reachable can create a client row; that is how the RFC works and what MCP clients need. The row grants nothing until a signed-in user approves it on the consent screen, and its redirect URIs are pinned from that moment. What is refused is anything that would make us an open redirector: exact-match redirect URIs only, https or loopback-http, no fragments.
Token audiences are bound to the host. aud is https://<host>/mcp, minted
from the same Host header core later hands to validateOAuthToken, so a token
obtained for one instance cannot be replayed against another.
openapi
Swagger UI at /openapi, the document at /openapi-json, covering all
223 routes the instance serves — 98 from this bundle, 125 from core.
API_DOCS=false turns it off.
Mounted outside /api deliberately: that prefix is guarded by a preHandler
hook in main.ts which 404s any request without a resolved workspace, and a
root-level path sidesteps it without editing the exclusion list. The SPA's *
catch-all does not shadow it — Fastify prefers an explicit route.
It mounted at /docs until upstream's public-spaces feature (876f3da1) gave
that path to PublicSpaceSeoController. Two GET handlers on one path is a
fatal FST_ERR_DUPLICATED_ROUTE at boot, not a compile error, so
upstream-contract/core-route-collisions.spec.ts now asserts our root mounts
stay clear of core's.
Schemas come from the @nestjs/swagger CLI plugin (patch 0010), which
reads TypeScript types, class-validator decorators and JSDoc at build time.
That is the whole reason this approach was viable: request bodies for core's
125 routes are described without annotating a single core file. @IsUUID()
becomes format: uuid, @IsIn([...]) becomes an enum, a doc comment becomes
the description.
Two things the generated document would get wrong, and how they are fixed:
- Responses are enveloped. Core's
TransformHttpResponseInterceptorwraps every payload in{ data, success, status }, so a naive spec is wrong for almost every route.applyResponseEnveloperewrites each 2xx schema to match. Handlers that write the reply themselves — SCIM, exports, attachments, health — are excluded by path prefix, since whether a handler takes@Res()is a source fact not visible at runtime. - Response shapes are not inferred. Nest cannot see a handler's return
type. Rather than emit a lie, operations get the envelope with an
undeclared
data. Annotating an EE handler with@ApiOkResponse({ type })improves it, and the envelope wrapper preserves what you declare.
The @nestjs/swagger dependency is the price. It is not in upstream's
package.json, so the lockfile moves too — and that ships as a whole file
(lockfile/pnpm-lock.yaml, substituted by the wrapper Dockerfile) rather than
as a patch hunk, because a diff against a generated, alphabetically ordered
lockfile breaks on nearly every upstream dependency bump. scripts/gen-lockfile.sh
regenerates it. It also drags in @scarf/scarf, whose postinstall script pnpm
11 refuses to run unapproved — which fails the Docker build until
pnpm-workspace.yaml says '@scarf/scarf': false. That is patch 0010 too.
To get the document as a file — for a code generator, or to diff what a change
did to the surface — scripts/gen-openapi.ts builds the same application and
writes it out. It needs a real environment, so run it in the container:
docker compose exec -w /app/apps/server docmost node dist/ee/scripts/gen-openapi.js > openapi.json
page-permission
Per-page access control: restrict a page, then grant users or groups
reader/writer. Restrictions inherit down the page tree.
Enforcement is core's, not ours. PagePermissionRepo already does the
recursive ancestor walk (canUserEditPage, getUserPageAccessLevel,
filterAccessiblePageIds), and core consumes it from PageAccessService,
favourites, labels, notifications and comment mentions. The tables
(page_access, page_permissions) ship with core too. Only the management
endpoints the client calls were missing — which is why /pages/permission-info
404'd while the licence advertised the feature.
This module is therefore a thin, guarded layer over the repo. Two guards worth knowing:
- Restricting grants the actor writer access. Otherwise the page ends up restricted with zero permission rows and the traversal check locks everyone out, including whoever just clicked the button.
- The last writer cannot be removed or demoted. A restricted page with only readers is unrecoverable through the UI — nobody can manage its permissions to add a writer back. The check runs inside a transaction and rolls back, rather than predicting the outcome up front.
Caching: canUserEditPage memoises per (userId, pageId) for 5s
(PERMISSION_CACHE_TTL_MS) and nothing in core invalidates it, so a change
can take that long to take effect. Invalidating properly would mean
enumerating every affected user across every descendant page.
template
Page templates: create one, edit it in a dedicated editor, then spawn pages from it. Templates are either global (workspace-wide) or scoped to a space.
Core owns the storage again — the templates table and TemplateRepo ship
with core, tsv is maintained by a database trigger off title +
text_content, and the allowMemberTemplates workspace setting is already
wired through workspace.service. Only the six endpoints were missing, which
is why the Templates button sat greyed out.
Access rules are taken from what the client already gates on, so the two agree (client gating is cosmetic; the service enforces):
| action | rule |
|---|---|
| view / use | global templates are workspace-wide; space ones need membership |
| create | workspace admin, or any member when allowMemberTemplates is on |
| make global | admin only |
| edit / delete | admin only |
/templates/use goes through PageService.create with the stored prosemirror
JSON, so the result is an ordinary collaborative page — correct slugId, tree
position and ydoc — rather than a special-cased copy. The target space is
independent of where the template lives, so a global template can seed a page
in any space the user belongs to.
personal-space
One private space per person, created by them. The "Allow personal spaces"
workspace toggle gates it (settings.spaces.allowPersonal).
Core owns everything underneath: spaces.is_personal with a partial unique
index on creator_id, so the database enforces one per person;
SpaceRepo.findPersonalSpace; and SpaceService.createSpace(..., {isPersonal}),
which creates the space, adds the creator as ADMIN and audits it — already
recording isPersonal in the audit changes. The toggle itself was already
licence-gated in workspace.service.
Only /personal-space/{info,create} were missing, which is why the toggle
claimed it needed an upgraded tier.
Two guards worth knowing:
- Duplicate check before insert. The unique index would reject a second personal space anyway, but as a constraint violation — a 500 carrying a Postgres message. Checking first returns something the caller can act on.
- Slug collisions are expected here.
SpaceService.createthrows on a duplicate slug rather than disambiguating, and personal space names collide readily ("Alex's space" twice). The service retries with a short suffix and falls back to the user id rather than looping.
group-role
A group can carry a workspace role, and everyone in it gets that role. The point is that admin rights follow Authentik group membership, so onboarding and offboarding happen in one place instead of two.
| mark | effect |
|---|---|
| Not assigned (default) | the group has no opinion; members' roles stay manual |
| Member | every member is held at member level |
| Admin | every member is a workspace admin |
| Owner | every member is promoted to owner — see below |
Set it on the group's detail page. Deliberately available on SCIM-synced groups, unlike everything else there: this mark is ours, not the IdP's, so no sync overwrites it. Marking the IdP's admins group as Admin is the whole point of the feature.
The receipt
Applying a grant also writes users.group_role. That column is what makes
withdrawal possible — without it, "this admin just lost their last admin
group" and "this admin was promoted by hand and is in no group at all" are
indistinguishable, and the first must be demoted while the second must not be
touched. It is also what greys out the role menu in the members table.
Highest grant wins, so someone in both a member group and an admin group is an admin; any other rule would depend on the order groups happen to be visited. Losing the last grant demotes to MEMBER rather than restoring whatever the person was before, because restoring would mean storing that too.
Owners are never managed
An Owner group promotes, and that is all it does: the promotion clears the receipt, so the user leaves group management for good and going back down is a manual act. Owner is the role that cannot be locked out and the one that repairs everything else, so no automatic process gets to take it away — not a group being unmarked, not an emptied claim, not an IdP that stopped answering.
The same reasoning runs through the rest of the guards:
- Only an owner can mark a group Owner, or change a group already marked
that way. Core refuses to let an admin act on the owner role directly
(
isAdminActingOnOwner), so they must not be able to do it sideways. - The default group cannot be marked at all — it contains everyone, so a grant on it would promote the entire workspace.
- The members table still offers Owner to owners, for a group-managed user, and the server allows exactly that one manual change. It is the escape hatch: whatever the IdP does, there is always a route to a role it cannot veto.
users.group_role has a CHECK admitting only member/admin, so "owner" is
not a state that table can even be in.
When it recomputes
Group membership is the input, so every path that changes it re-derives:
| trigger | where |
|---|---|
| the mark itself changes | POST /groups/role (ours) |
| SCIM membership sync, group delete | scim/services/scim-group.service.ts (ours) |
| OIDC group claim, on every login | sso/services/sso.service.ts (ours) |
| a member is added or removed by hand | core group-user.service.ts (patch 0009) |
| a group is deleted | core group.service.ts (patch 0009) |
The core triggers go through the usual require() + ModuleRef hook, so
without this bundle they are no-ops and roles stay entirely manual — upstream's
behaviour. All of them skip unmarked groups, which is nearly all groups.
Recomputes never throw at their caller. Each one is a side effect of something else the user asked for, and failing to re-derive a role must not fail a login.
OIDC group sync
Turn on "Group sync" on the OIDC provider and the groups claim is
reconciled into Docmost groups on every login, not just the first.
Groups are matched by name and created on demand, so a group SCIM has not
provisioned yet still works — the first person to log in carrying it brings it
into existence. Created groups are marked is_external, the same mark SCIM
uses, so patch 0008 locks them in the UI.
Requesting the claim is what makes this possible: the authorization request
adds the groups scope when group sync is on, and if the ID token still lacks
it we fall back to the userinfo endpoint. Both only happen when group sync is
on, so a provider without it costs no extra round trip and cannot be broken by
an IdP that does not know the scope.
Removal is governed by SSO_OIDC_GROUP_SYNC:
| value | behaviour |
|---|---|
add (default) |
join groups named in the claim, never leave any |
reconcile |
also leave IdP-managed groups the claim no longer names |
add is the default because it cannot fight SCIM. SCIM reconciles a group to
exactly the members the IdP sent, while an OIDC claim is per-user and may
legitimately omit groups SCIM manages — under strict reconcile each side would
keep undoing the other on alternating login and sync. Use reconcile when
OIDC is your only group source.
Either way only is_external groups are ever left, and the default group is
never touched: a manual assignment must not be undone by a login. Failures are
logged and swallowed — group sync is a convenience, and a transient database
error during it must not become a failed login.