MCP Hub
Back to servers

desktop-touch-mcp

Windows desktop automation MCP server — screenshot, mouse, keyboard & UI Automation. Lets LLM agents see and control your Windows desktop directly.

glama
Updated
Apr 13, 2026

desktop-touch-mcp

日本語

Glama Link

Stop pasting screenshots. Let Claude see and control your desktop directly.

An MCP server that gives Claude eyes and hands on Windows — 25 tools covering screenshots, mouse, keyboard, and Windows UI Automation, designed from the ground up for LLM efficiency.

Applies MPEG P-frame diffing to window capture: only changed windows are sent after the first frame, cutting token usage by ~60–80% in typical automation loops.


Features

  • LLM-native design — Built around how LLMs think, not how humans click. run_macro batches multiple operations into a single API call; diffMode sends only the windows that changed since the last frame. Minimal tokens, minimal round-trips.
  • Full CJK support — Uses Win32 GetWindowTextW for window titles, avoiding nut-js garbling. IME bypass input supported for Japanese/Chinese/Korean environments.
  • 3-tier token reductiondetail="image" (~443 tok) / detail="text" (~100–300 tok) / diffMode=true (~160 tok). Send pixels only when you actually need to see them.
  • 1:1 coordinate modedotByDot=true captures at native resolution (WebP). Image pixel = screen coordinate — no scale math needed. With origin+scale passed to mouse_click, the server converts coords for you — eliminating off-by-one / scale bugs.
  • Chrome / AWS data reductiongrayscale=true (~50% size), dotByDotMaxDimension=1280 (auto-scaled with coord preservation), windowTitle + region sub-crop (exclude browser chrome). Targeted at heavy dotByDot captures. Typical token reduction: 50–70%.
  • Chromium smart fallbackdetail="text" on Chrome/Edge/Brave auto-skips UIA (prohibitively slow there) and runs Windows OCR. hints.chromiumGuard + hints.ocrFallbackFired flag the path taken.
  • UIA element extractiondetail="text" returns button names and clickAt coords as JSON. Claude can click the right element without ever looking at a screenshot.
  • Auto-dock CLIdock_window snaps any window to a screen corner with always-on-top. Set DESKTOP_TOUCH_DOCK_TITLE='@parent' to auto-dock the terminal hosting Claude on MCP startup — the process-tree walker finds the right window regardless of title.
  • Emergency stop (Failsafe) — Move the mouse to the top-left corner (within 10px of 0,0) to immediately terminate the MCP server.

Requirements

OSWindows 10 / 11 (64-bit)
Node.jsv20+ recommended (tested on v22+)
PowerShell5.1+ (bundled with Windows)
Claude CLIclaude command must be available

Note: nut-js native bindings require the Visual C++ Redistributable. Download from Microsoft if not already installed.


Installation

git clone https://github.com/Harusame64/desktop-touch-mcp.git
cd desktop-touch-mcp
npm install
npm run build

Register with Claude CLI

Add to ~/.claude.json under mcpServers:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"]
    }
  }
}

Note: Replace D:/path/to/desktop-touch-mcp with the actual path where you cloned this repository.

No system prompt needed. The command reference is automatically injected into Claude via the MCP initialize response's instructions field.


Tools (34 total)

Screenshot (5)

ToolDescription
screenshotMain capture. Supports detail, dotByDot, dotByDotMaxDimension, grayscale, region sub-crop, diffMode
screenshot_backgroundCapture a background window without focusing it (PrintWindow API)
screenshot_ocrWindows.Media.Ocr on a window; returns word-level text + screen clickAt coords
get_screen_infoMonitor layout, DPI, cursor position
scroll_captureFull-page stitch by scrolling (MAE overlap detection + 10% fallback)

Window management (4)

ToolDescription
get_windowsList all windows in Z-order
get_active_windowInfo about the focused window
focus_windowBring a window to foreground by partial title match
dock_windowSnap a window to a screen corner at a small size + always-on-top (for keeping CLI visible)

Mouse (5)

ToolDescription
mouse_move / mouse_click / mouse_dragMove, click, drag. Accept speed and homing parameters
scrollScroll in any direction. Accepts speed and homing parameters
get_cursor_positionCurrent cursor coordinates

Keyboard (2)

ToolDescription
keyboard_typeType text (use_clipboard=true bypasses IME)
keyboard_pressKey combos (ctrl+c, alt+f4, etc.)

UI Automation (4)

ToolDescription
get_ui_elementsFull UIA element tree for a window
click_elementClick a button by name or automationId — no coordinates needed
set_element_valueWrite directly to a text field
scope_elementHigh-res zoom crop of an element + its child tree

Browser CDP (7)

ToolDescription
browser_connectConnect to Chrome/Edge via CDP; lists open tabs
browser_find_elementCSS selector → exact physical screen coords
browser_click_elementFind DOM element + click in one step
browser_evalEvaluate JS expression in the browser tab
browser_get_domGet outerHTML of element or document.body
browser_navigateNavigate via CDP Page.navigate (no address bar needed)
browser_disconnectClose cached CDP WebSocket sessions

Workspace (2)

ToolDescription
workspace_snapshotAll windows: thumbnails + UI summaries in one call
workspace_launchLaunch an app and auto-detect the new window

Pin / Macro (3)

ToolDescription
pin_window / unpin_windowAlways-on-top toggle
run_macroExecute up to 50 steps sequentially in one MCP call

Browser CDP automation

For web automation, connect Chrome or Edge with the remote debugging port enabled — no Selenium or Playwright needed.

# Launch Chrome in CDP mode
chrome.exe --remote-debugging-port=9222 --user-data-dir=C:\tmp\cdp
browser_connect()                       → list open tabs + get tabIds
browser_find_element("#submit")         → CSS selector → physical screen coords
browser_click_element("#submit")        → find + click in one step (auto-focuses browser)
browser_eval("document.title")          → evaluate JS, returns result
browser_get_dom("#main", maxLength=5000)→ outerHTML, truncated to maxLength chars
browser_navigate("https://example.com") → navigate via CDP (no address bar interaction)
browser_disconnect()                    → clean up WebSocket sessions

Coordinates returned by browser_find_element account for the browser chrome (tab strip + address bar height) and devicePixelRatio, so they can be passed directly to mouse_click without any scaling.

Recommended web workflow:

browser_connect() → browser_get_dom() → browser_find_element(selector) → browser_click_element(selector)

Auto-dock CLI on startup

Keep Claude CLI visible while operating other apps full-screen. Set env vars in your MCP config and the docked window auto-snaps into place every MCP startup.

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"],
      "env": {
        "DESKTOP_TOUCH_DOCK_TITLE": "@parent",
        "DESKTOP_TOUCH_DOCK_CORNER": "bottom-right",
        "DESKTOP_TOUCH_DOCK_WIDTH": "480",
        "DESKTOP_TOUCH_DOCK_HEIGHT": "360",
        "DESKTOP_TOUCH_DOCK_PIN": "true"
      }
    }
  }
}
Env varDefaultNotes
DESKTOP_TOUCH_DOCK_TITLE(unset = off)@parent walks the MCP process tree to find the hosting terminal — immune to title / branch / project changes. Or use a literal substring.
DESKTOP_TOUCH_DOCK_CORNERbottom-righttop-left / top-right / bottom-left / bottom-right
DESKTOP_TOUCH_DOCK_WIDTH / HEIGHT480 / 360px ("480") or ratio of work area ("25%") — 4K/8K auto-adapts
DESKTOP_TOUCH_DOCK_PINtrueAlways-on-top toggle
DESKTOP_TOUCH_DOCK_MONITORprimaryMonitor id from get_screen_info
DESKTOP_TOUCH_DOCK_SCALE_DPIfalseIf true, multiply px values by dpi / 96 (opt-in per-monitor scaling)
DESKTOP_TOUCH_DOCK_MARGIN8Screen-edge padding (px)
DESKTOP_TOUCH_DOCK_TIMEOUT_MS5000Max wait for the target window to appear

Input routing gotcha: when a pinned window is active (e.g. Claude CLI), keyboard_type / keyboard_press send keys to it, not the app you wanted to type into. Always call focus_window(title=...) before keyboard operations, then verify isActive=true via screenshot(detail='meta').


Mouse homing correction

When Claude calls screenshot(detail='text') to read coordinates and then mouse_click seconds later, the target window may have moved. The homing system corrects this automatically.

TierHow to enableLatencyWhat it does
1Always-on (if cache exists)<1msApplies (dx, dy) offset when window moved
2Pass windowTitle hint~100msAuto-focuses window if it went behind another
3Pass elementName/elementId + windowTitle1–3sUIA re-query for fresh coords on resize
# Tier 1 only (automatic)
mouse_click(x=500, y=300)

# Tier 1 + 2: also bring window to front if hidden
mouse_click(x=500, y=300, windowTitle="メモ帳")

# Tier 1 + 2 + 3: also re-query UIA if window resized
mouse_click(x=500, y=300, windowTitle="メモ帳", elementName="保存")

# Traction control OFF — no correction
mouse_click(x=500, y=300, homing=false)

The homing parameter is available on mouse_click, mouse_move, mouse_drag, and scroll. The cache is updated automatically on every screenshot(), get_windows(), focus_window(), and workspace_snapshot() call.

mouse_click image-local coords (origin + scale)

When you take a dotByDot screenshot with dotByDotMaxDimension, the response prints the origin and scale values. Instead of computing screen coords manually, copy them into mouse_click:

# Screenshot response:
#   origin: (0, 120) | scale: 0.6667
#   To click image pixel (ix, iy): mouse_click(x=ix, y=iy, origin={x:0, y:120}, scale=0.6667)

mouse_click(x=640, y=300, origin={x:0, y:120}, scale=0.6667, windowTitle="Chrome")
# Server converts: screen = (0 + 640/0.6667, 120 + 300/0.6667) = (960, 570)

This eliminates a whole class of off-by-one and scale bugs. Without origin/scale, x/y remain absolute screen pixels (unchanged behavior).


screenshot key parameters

detail="image"          — PNG/WebP pixels (default)
detail="text"           — UIA element JSON + clickAt coords (no image, ~100–300 tok)
detail="meta"           — Title + region only (cheapest, ~20 tok/window)
dotByDot=true           — 1:1 WebP; image_px + origin = screen_px
dotByDotMaxDimension=N  — cap longest edge (response includes scale for coord math)
grayscale=true          — ~50% smaller for text-heavy captures (code/AWS console)
region={x,y,w,h}        — with windowTitle: window-local coords (exclude browser chrome)
                          without: virtual screen coords
diffMode=true           — I-frame first call, P-frame (changed windows only) after (~160 tok)
ocrFallback="auto"      — detail='text' auto-fires Windows OCR on uiaSparse or empty

Recommended Chrome combo (50–70% data reduction):

screenshot(windowTitle="Chrome",
           dotByDot=true, dotByDotMaxDimension=1280, grayscale=true,
           region={x:0, y:120, width:1920, height:900})  # skip browser chrome

Recommended workflow:

workspace_snapshot()                     → full orientation (resets diff buffer)
screenshot(detail="text", windowTitle=X) → get actionable[].clickAt coords
mouse_click(x, y)                        → click directly, no math needed
screenshot(diffMode=true)                → check only what changed (~160 tok)

Security

Emergency stop (Failsafe)

Move the mouse to the top-left corner of the screen (within 10px of 0,0) to immediately terminate the MCP server.

  • Per-tool check: checkFailsafe() runs before every tool handler
  • Background monitor: 500ms polling as a backup for long-running operations
  • Trigger radius: 10px

Blocked operations

workspace_launch blocklist: cmd.exe, powershell.exe, pwsh.exe, wscript.exe, cscript.exe, mshta.exe, regsvr32.exe, rundll32.exe, msiexec.exe, bash.exe, wsl.exe are blocked. Script extensions (.bat, .ps1, .vbs, etc.) are rejected. Arguments containing ;, &, |, `, $(, ${ are also rejected.

keyboard_press blocklist: Win+R (Run dialog), Win+X (admin menu), Win+S (search), Win+L (lock screen) are blocked.

PowerShell injection protection

All -like patterns in the UIA bridge are sanitized with escapeLike(), which escapes wildcard characters (*, ?, [, ]) before they reach PowerShell.

Allowlist for workspace_launch

Shell interpreters are blocked by default. To allow specific executables, create an allowlist file:

File locations (searched in order):

  1. Path in DESKTOP_TOUCH_ALLOWLIST environment variable
  2. ~/.claude/desktop-touch-allowlist.json
  3. desktop-touch-allowlist.json in the server's working directory

Format:

{
  "allowedExecutables": [
    "pwsh.exe",
    "C:\\Tools\\myapp.exe"
  ]
}

Changes take effect immediately — no restart needed.


Mouse movement speed

All mouse tools (mouse_move, mouse_click, mouse_drag, scroll) accept an optional speed parameter:

ValueBehavior
OmittedUses the configured default (see below)
0Instant teleport — setPosition(), no animation
1–NAnimated movement at N px/sec

Default speed is 1500 px/sec. Change it permanently via the DESKTOP_TOUCH_MOUSE_SPEED environment variable:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"],
      "env": {
        "DESKTOP_TOUCH_MOUSE_SPEED": "3000"
      }
    }
  }
}

Common values: 0 = teleport, 1500 = default gentle, 3000 = fast, 5000 = very fast.


Known limitations

LimitationDetailWorkaround
Games / video players may return black or hang in background captureDirectX fullscreen apps may not work even with PW_RENDERFULLCONTENTRetry with screenshot_background(fullContent=false); if still black, use foreground screenshot
UIA call overhead~300ms per call via PowerShell; workspace_snapshot uses a 2s timeout internallyBatch with workspace_snapshot upfront, then use diffMode for incremental checks
Chrome / WinUI3 UIA elements are emptyChromium exposes only limited UIAscreenshot(detail='text') auto-detects Chromium and falls back to Windows OCR (hints.chromiumGuard=true). For richer DOM access use browser_connect + browser_find_element
Chromium title-regex misses when sites rewrite document.titleGuard relies on the - Google Chrome suffix being present; some sites push it off the end of a long titleTitle is treated as plain Chrome (UIA runs). OCR path is still reachable via ocrFallback='always' or when UIA returns <5 elements (uiaSparse)
browser_* CDP tools need Chrome launched with --remote-debugging-portIf Chrome is already running on the default profile without the flag, browser_launch / browser_connect fail. The CDP E2E suite (tests/e2e/browser-cdp.test.ts) will also fail in that stateClose Chrome first, then browser_launch will relaunch it in debug mode, or start Chrome manually with --remote-debugging-port=9222 --user-data-dir=C:\tmp\cdp
Layer buffer TTLBuffer auto-clears after 90s of inactivity → next diffMode becomes an I-frameAfter long waits, call workspace_snapshot to explicitly reset the buffer
keyboard_type / keyboard_press follow focusWhen dock_window(pin=true) keeps another window on top (e.g. Claude CLI), keystrokes may be absorbed by that windowCall focus_window(title=...) first and verify isActive=true via screenshot(detail='meta') before sending keys

Token cost reference

ModeTokensUse case
screenshot (768px PNG)~443 tokGeneral visual check
screenshot(dotByDot=true) window~800 tokPrecise clicking (no coordinate math)
screenshot(diffMode=true)~160 tokPost-action diff
screenshot(detail="text")~100–300 tokUI interaction (no image)
workspace_snapshot~2000 tokFull session orientation

License

MIT

Reviews

No reviews yet

Sign in to write a review