Skip to content
Median
Esc
↑↓navigate↵open⌘Jpreview
On this page

Connected sources

Import Notion pages, GitHub repositories, docs sites and websites, and keep them current.

Open Knowledge, press Add, then Add content. The Sources list has Notion, GitHub, Docs site and Website.

Only admins and owners can open a source. Everyone else sees Admin access required. See roles.

Refresh and limits

Source Refreshes on its own By hand Limits
Notion When the agent starts a reply, at most every 15 minutes. Only pages already imported Library sync button 20 pages per import. Sub-pages 3 levels deep, 25 per page
GitHub repository On every push to the synced branch. When the agent starts a reply, at most once a minute per repository Sync now, library sync button 500 markdown files. 900 KB per file
Docs site Same as a repository Sync now, library sync button 500 pages
Website Never Re-scrape 500 pages per crawl. 3 crawls at once
  • A refresh runs beside the reply. The reply uses what is indexed at that moment, and the next reply gets the changes.
  • The library sync button has the tooltip Pull fresh copies from GitHub and Notion, or names just the one the library holds. It reads every file in every repository and docs site again, and pulls every Notion page again while Notion is connected.
  • Sync now also reads every file, not just the ones that changed. Pressed during a sync, either button runs once that sync finishes, and so does a push.
  • Titles come from a file’s frontmatter title, then its first heading, then its filename. It appears once the library holds a GitHub document, or a Notion page with Notion still connected.
  • A repository or docs site whose last sync failed shows up under Needs fixing on the Dashboard as owner/repo is not syncing.

Notion

Connect

Open Notion and press Connect Notion. Notion’s consent screen opens. Choose the pages Median may see there.

Pick pages

Under Pages, tick pages in the list, or type in Search pages, or paste a link. The list shows 20 pages at a time, most recently edited first. A pasted link imports that one page.

Import

Leave Include sub-pages on to bring child pages along. Press Import. The library opens and each page arrives with an Indexing badge.

Rule Value
Pages per import 20. More returns “Up to 20 pages at a time.”
Sub-page depth 3 levels below each picked page
Sub-pages per page 25
Sub-page ceiling New sub-pages stop once the organization holds 300 Notion documents
Title The page’s name in Notion
Editing Edit in Notion. The document page shows a Read only badge with the tooltip “Edit this document in Notion.”

A refresh pulls every imported page again. It does not add new pages or new sub-pages. Import those.

Choose pages on the Workspace row reopens Notion’s consent screen, where you can share more pages with Median.

Disconnect asks Disconnect Notion? and keeps every imported document. Refreshes stop.

Connect error Cause
Notion connection was not authorized. Try connecting again. Consent was declined on Notion
The connection did not finish. Try again. Notion did not complete the handoff

A connect link works once, for 15 minutes. After that, the return from Notion lands on a page that reads “This connection link has expired or was already used. Start again from Median.”

Error on a page Cause
This page has nothing on it yet. The page is empty
This page is more than one document can hold. Split it up. The page converts to more than 900 KB of markdown
Notion is not connected anymore. Connect it, then try again. The workspace was disconnected
The Notion connection expired. Connect it again, then retry. Notion would not renew access
That does not look like a link to a Notion page. The pasted link is not a Notion page

A failed page shows Failed in the library. The reason and Try again are on the document page.

GitHub

Median reads repositories through its GitHub app. You pick which repositories the app can see during the install on GitHub.

Connect an account

Open GitHub and press Connect GitHub. Install the app on GitHub and pick repositories. You come back to the GitHub page.

Add a repository

Press Add repository, find it with Search repositories, and press Import.

Set the fields

Pick a Branch and a Folder, fill Published at if you want links, and press Add. The first sync starts at once.

Field Default Notes
Branch The repository’s default branch The default branch is followed even if it is renamed later
Folder Sync the whole repository Each folder shows how many markdown and spec files it holds
Published at Empty Optional. Where the files are published, so the agent can link them. See Published links
  • Connect another under Accounts adds a second GitHub account or organization. The account menu in the repository picker has Connect another account too.
  • The picker lists recently active repositories. Type a full owner/repo to reach one the list does not show.
  • Missing a repository? Add access on GitHub opens the app’s settings on GitHub.
  • One repository can be added more than once with different folders. The same repository and folder twice returns “That repository is already connected.”

What a sync takes

File Becomes
.md, .mdx, .markdown under the folder One document per file. The title is the first heading, or the file name
An OpenAPI or Swagger spec One document per endpoint, plus an overview. See API specs
  • A sync compares every file with the last sync. Unchanged files are skipped. Changed files are indexed again.
  • A file deleted from the repository leaves the library.
  • A new file joins the library folder its neighbours in the repository are in.
  • Markdown files over 900 KB are skipped.
  • Documents whose indexing failed are retried on every sync.
  • Synced documents are edited in the repository. The document page shows a Read only badge with the tooltip “Edit this document in its source repository.”

Repository rows

Each row shows owner/repo, the folder and branch when set, and a status line.

Status Means
Syncing now A sync is running, including the first one after you add it
Synced 5m ago The last sync finished
The last sync failed The error shows under the row
Button Does
Published docs Opens the address dialog. See Published links
Sync now Starts a sync. Does nothing while one is running
Remove Asks Remove owner/repo?, then removes the repository and its documents

Disconnect on an account removes every repository and docs site read through it, and their documents. Uninstalling the app on GitHub, or taking a repository’s access away there, does the same.

Sync errors

Error Fix
More than 500 markdown files in there. Point the sync at a docs folder. Remove the repository and add it again with a folder
This docs site has more than 500 pages, which is more than a sync can take. The site grew past 500 pages. Remove it and add the repository as a plain repository with a smaller folder
This repository is too large to sync. Select a documentation folder instead. GitHub returned a partial file tree for the branch. The whole branch is read even when a folder is set, so a folder does not clear this
GitHub could not find that. Check the repository and branch, and that the app can see it. The branch is gone, or the app lost access
GitHub turned us away. Check the app’s key, and that it is still installed. The app was uninstalled or lost permission
GitHub is not answering right now. Try again in a bit. GitHub was down. Press Sync now later
Could not sync this repository. Try again. Anything else

A sync that runs longer than 10 minutes is treated as stuck. The next sync takes over.

Connect error Cause
This connection attempt expired. Connect again. More than an hour passed on GitHub
That GitHub installation is connected to another organization. Disconnect it there first. One GitHub install belongs to one Median organization

API specs

Repositories and docs sites turn OpenAPI specs into documents. Each endpoint document has Parameters, Request body, Responses and a curl command under Example request. The overview document lists the version, Base URL, Authentication and every endpoint.

Each spec can also go on your site as an API reference.

A file is read as a spec when all three tests pass:

Test Passes
Extension .json, .yaml or .yml
Name or folder The name contains openapi or swagger. Or the name has api as its own word, bounded by the start or end or by -, _ or ., so api.json and public-api.yaml pass and rapid.json does not. Or a parent folder is named openapi, swagger, spec, specs or apis
Contents It parses, has an openapi or swagger version key, and has a paths object

A file that fails the contents test is ignored.

Limit Value
Spec files read per sync 25
Spec file size 5 MB
Endpoints per spec 300. The rest are left out
Versions OpenAPI 3.x and Swagger 2.0

Specs must sit under the synced folder. A docs site also takes specs its config names, wherever they sit.

Docs sites

A docs site is a GitHub repository read through its framework’s config. The config sets which files are pages and the address each page publishes at. It uses the same GitHub accounts as the GitHub source.

Open Docs site and press Add docs site. The dialog has three steps.

Step What you do
Repository Pick the repository. Continue
Framework Check the Branch. Median reads the branch and shows the Framework with its config file, and Pages with any API specs. Continue
Address Fill Published at. It is required. Three sample files show the address each will publish at, and update as you type. Add docs site

A Blume or Docusaurus config that names its site fills Published at for you.

Framework Config Pages and routes
Mintlify docs.json or mint.json Only pages in the navigation, at the route the navigation gives. Specs from any openapi field
Docusaurus docusaurus.config.js, .ts, .mjs, .cjs or .mts Files under docs.path, default docs. Published under routeBasePath, default docs. Number prefixes like 01- are dropped. url plus baseUrl fill the address. Specs from specPath
Fumadocs source.config.ts, .js, .mjs or .mts Files under defineDocs({ dir }), default content/docs. Published under the folder holding the [[...slug]] route, default docs
GitBook .gitbook.yaml or SUMMARY.md Only pages the summary links. root and structure.summary are read from the yaml
Blume blume.config.ts, .js, .mjs, .mts or .json Files under content.root, default docs. deployment.site fills the address. Specs from openapi when enabled: true
  • index and README files publish at their folder’s route, except on Mintlify.
  • A route prefix is not added twice when Published at already ends with it.
  • When a repository holds several configs, GitBook is tried last.
  • Median reads configs as text and does not run them. Check the three sample addresses before you add the site.
Situation What you see
No config found Not recognised, “Median reads Mintlify, Docusaurus, Fumadocs, GitBook and Blume.” and a Sync it as a plain repository link to the GitHub source
Over 500 pages “This site has more than 500 pages, which is more than a sync can take.” Continue stays disabled. Add it as a plain repository with a smaller folder instead
Config names more than 100 API specs “This site exceeds the page limit for a single sync.”

The site row shows the framework, the content folder and the branch, with the same status line and buttons as a repository row. Sync now on a docs site reads the config again even when the branch has not changed. Unchanged files are still skipped. Every sync that finds the repository changed reads the config again too, so a moved docs folder follows.

Websites

Websites need a paid plan. Without one the page shows “Available on Standard and Pro.” See plans.

Paste an address

Open Website and type the address, for example docs.example.com. https:// is added when missing.

Scrape

Press Scrape. The crawl appears under Sites, and pages land in the library as they are scraped.

  • The crawl stays under the address you give. example.com/docs does not read example.com/blog.
  • It reads the site’s sitemap and follows links. PDFs are read too.
  • Each scraped page is charged in credits. A crawl stops at 500 pages, or sooner when credits run short.
  • Each page becomes a document linked to the address it was scraped from.

A running crawl opens into three stages.

Stage Shows
Discover Pages found
Scrape Pages scraped, as 12 of 40
Index Pages indexed. Starts once scraping ends
Crawl Row reads
Running “Started 2m ago by” and the name of whoever started it
Finished “Scraped 5m ago by” and the same name. The page count, a skipped count for pages that came back blank or too large, and a kept count for pages you edited sit on the right
Failed The error, with the same counts on the right
Button When Does
Cancel Discover or Scrape running Stops the crawl. Pages already scraped stay and are indexed
Dismiss After scraping ends Removes the row. The pages stay in the library
Re-scrape Crawl finished or failed Fetches live copies and updates pages in place. Pages you edited keep your version
  • Nothing refreshes on a schedule. Press Re-scrape when the site changes.
  • A first crawl may use copies up to 2 days old. A re-scrape, or a new crawl of an address you scraped before, fetches live pages.
  • The list shows the 20 most recent crawls.
Error Cause
That does not look like a site address. Not a web address
That site is already being scraped. A crawl of that address is running
Up to 3 sites at a time. Let one finish first. 3 crawls are running
Nothing readable came back from this site. No page had content
The crawl took too long, so it was let go. Try again. The crawl ran for 60 minutes
The page import service is busy. Try again in a minute. The scraper was busy
The scraped pages could not be indexed. Try the site again. Indexing failed after the scrape
Your organization is out of credits. Add more in Settings, Billing. No credits were left when the crawl started

Other failures say to try again later.

From the API

Request Does
POST /v1/knowledge/sync Same as the library sync button
POST /v1/integrations/repos Adds a repository with repo, branch, path and publishedAt
PATCH /v1/integrations/repos/{owner}/{repo} Sets Published at
DELETE /v1/integrations/repos/{owner}/{repo} Stops syncing a repository. Its documents stay in the library, unlike Remove in the dashboard
GET and POST /v1/knowledge/crawls Lists crawls or starts one with url
DELETE /v1/knowledge/crawls/{id} Cancels or dismisses a crawl

Full schemas are in the management API reference.

Was this page helpful?