Vidoc Security Lab vidocsecurity.com

Built on GLM-5.3 · Mythos-level cyber capability

The most cyber‑capable open model, for defenders

vidoc-1 runs GLM-5.3, the open model that independent evaluators rank at the frontier of cyber capability. We put it in the hands of security teams for application security, authorized penetration testing, reverse engineering and research. Point Claude Code, Codex or opencode at it by changing one URL. No sign-up and no API key.

base_url

Need unlimited or compliance-ready access? Talk to us.

"The most cyber-capable open-weight model released to date."
NIST CAISI assessment of GLM-5.3, Sep 2026
Builds working exploits end to end, on par with Claude Mythos Preview.
Anthropic Frontier Red Team, GLM-5.3 analysis
100,000,000 tokens left

Every IP gets 100 million tokens. The squares run from a full allowance to an empty one; the outlined square is where yours is now.

Previous-generation models scored zero

Anthropic's Frontier Red Team measured whether models can find and exploit memory-safety bugs in open-source software and land a full control-flow hijack. GLM-5.3 crosses a line that last generation's models never reached.

binary-exploitation.bench
Claude Mythos PreviewAnthropic · limited release
6%
GLM-5.3vidoc-1 runs this
4%
Claude Opus 4.6previous generation
0%
GLM-5.2previous generation
0%
0%2%4%6%8%

Full control-flow hijacks over 100 OSS-Fuzz tasks. On V8 end-to-end exploits (ExploitBench), GLM-5.3 succeeds 50 times in 410 attempts, close to Claude Mythos Preview's 56. Source: Anthropic Frontier Red Team.

Security work it handles

Code review, vulnerability triage, authorized penetration tests, reverse engineering, malware analysis, threat modeling, CTF practice and report writing. A few examples of what a session looks like:

Application security

orders.py
> review this handler for vulnerabilitiesline 14 builds SQL from the URL path, soorder_id can change the query (SQL injection).- cur.execute(f"... WHERE id = {order_id}")+ cur.execute("... WHERE id = %s", (order_id,))▌

Penetration testing

staging-scan.txt
> rank what to check first on our staging host1  8080/tcp  admin panel over plain HTTP   confirm it should be reachable at all2  22/tcp    OpenSSH 7.4, years out of date   check vendor advisories, plan an upgrade3  443/tcp   TLS 1.0 still enabled   disable old protocol versions▌

Reverse engineering

sub_401a20.c
> what does this function do?It XORs each byte of a buffer with 0x5A andreturns the result: a string decoded at runtime, a common way to hide configuration inmalware. The decoded value is an update URL.renamed  sub_401a20 → decode_config_string▌

Security research

threat-model.md
> threat model our password reset flowhigh    reset tokens never expiremedium  token travels in a GET query string        and ends up in proxy logsmedium  no rate limit on /resetlow     reset email reveals if an account exists▌

These are illustrations of the kinds of tasks people use it for. Use it only on systems, code and binaries you own or are authorized to test.

Sandbox: capture the flag

An isolated target, a bug someone left in, and a flag to find. Type help to begin. None of this touches a real system — it's a playground.

vidoc-sandbox — target acme.local (isolated) unsolved

A deliberately vulnerable mock, for practice. The one rule of this work: only test what you're authorized to — which here is everything, because none of it is real.

Set up your agent

vidoc-1 speaks the OpenAI Chat Completions, OpenAI Responses and Anthropic Messages APIs, so each client only needs a new base URL and model name.


    

The model

What your agent gets on the other end of the URL.

vidoc-1.spec
base model
GLM-5.3 (Zhipu AI / Z.ai), the open model NIST CAISI rated the most cyber-capable open-weight release to date.
context
131,072 tokens, shared by the prompt and the output.
thinking
Always on, at low effort by default to keep responses fast. Pass reasoning_effort: "high" for deeper reasoning.
tool calling
Supported in all three APIs.
endpoints
/v1/chat/completions, /v1/responses, /v1/messages, /v1/messages/count_tokens and /v1/models
api key
None. Clients that require one can send any value.
model name
vidoc-1. Requests that name another model are served by vidoc-1 too.

Limits

  • Authorized use only. Test only systems, code and binaries you own or have permission to test.
  • 100 million tokens per IP address, counting prompt and output tokens. Prompt tokens served from the model's cache count half.
  • Coding agents resend the whole conversation on every turn, so a long session uses tokens quickly.
  • When your IP runs out, requests return HTTP 429.
  • A stream you cancel is charged an estimate of what it used.
  • Request bodies can be at most 4.5 MB.
  • Requests and responses are stored for security and compliance purposes. Don't send secrets or code you aren't allowed to share.
  • The model sleeps when nobody uses it. The first request after a quiet period waits while it starts, which can take several minutes. If it times out, try again a little later.
for-teams.txt

Need more than the free tier?

For unlimited access, or access that meets your organisation's compliance requirements, talk to us and we'll set it up for your team.