bimq

bimq

Read-only BIM query server for agents: structured, policy-bounded queries on IFC/gbXML models with deterministic cited results.

Category
访问服务器

README

bimq

A read-only BIM query server for agents. Point it at an IFC or gbXML model and it answers structured questions — fire-rated doors on level 3, elements with no material assigned, spaces below the minimum daylight area — bounded by a policy file, with deterministic results and a citation back to the GlobalId and source line behind every row.

Zero dependencies. Python 3.11+. MCP server over stdio, plus a CLI that answers the same questions so you can check a policy before you trust an agent to it.

bimq query elements_by_property model.ifc \
    type=IfcDoor storey="Level 3" property=FireRating op=exists
id                      type     name      tag   storey_name  source          match
----------------------  -------  --------  ----  -----------  --------------  -------------------------------
0XBbD$nZDLuRru91_CQ_xe  IfcDoor  Door-302  D302  Level 3      office.ifc:314  Pset_DoorCommon.FireRating=EI60
31kamnSrrNbf3eF0_vhXrJ  IfcDoor  Door-301  D301  Level 3      office.ifc:304  Pset_DoorCommon.FireRating=EI60

2 row(s)
digest: sha256:06321f331735417dd149d649b8e26de71a63ff2bb26a31c38cbf66a4f0314b77

Then check it, because a citation you cannot follow is just a confident-looking string:

bimq cite model.ifc 31kamnSrrNbf3eF0_vhXrJ
31kamnSrrNbf3eF0_vhXrJ  (IfcGloballyUniqueId)
office.ifc:304  #297

#297= IFCDOOR('31kamnSrrNbf3eF0_vhXrJ',#5,'Door-301',$,$,$,$,'D301',2100.0,900.0,.DOOR.,.SINGLE_SWING_LEFT.,$);

Why

The current instinct is to dump IFC text into a context window. That fails immediately at real model sizes, and it fails quietly: a 300 MB model is roughly 95% geometry, so what fits in the window is a truncated arbitrary slice, and the model answers from it anyway. The failure looks like a fluent paragraph about a door that does not exist.

bimq inverts it. The model stays on disk. Queries are structured, the answers are small, and every row carries the id and line it came from — so a claim can be checked against the file instead of trusted.

Three properties hold for every answer:

Bounded. A TOML policy file says what is readable — which files, which queries, which types, which storeys, which properties. The engine reduces the model to the visible set before the query runs, so a query primitive cannot reach what the policy hides even by accident.

Deterministic. Same model, same query, same bytes. Every answer carries a digest you can pin in a test. Element order, group order and float rounding are all fixed; the read block size and the file's name do not change a finding.

Cited. Every row carries {id, id_kind, source, ref, line}. bimq cite resolves it back to the original statement. The test-suite re-reads the recorded line for every element of every fixture and fails if the id is not there.

Install

pip install bimq

Or run it from a clone with no install at all — there is nothing to build:

python -m bimq describe tests/fixtures/office.ifc

Use it as an MCP server

{
  "mcpServers": {
    "bimq": {
      "command": "bimq",
      "args": ["serve", "/srv/bim/tower.ifc", "-p", "/srv/bim/policy.toml"]
    }
  }
}

Every query primitive becomes a bim_* tool, all annotated readOnlyHint, plus bim_cite. Omit the model path to let each call name its own file — then allow_sources is what stands between a path argument and your filesystem.

The server tells the agent how to behave on initialize: call bim_model_summary first, quote a GlobalId for anything you assert, treat truncated as "there are more", and read notes — because no results and no data recorded are different findings and only the notes distinguish them.

A policy refusal comes back as a successful tool result carrying policy_denied and the rule that fired, not as a protocol error. An agent that receives a protocol error retries; an agent told "this policy does not expose costs" reports the limit and moves on.

Query primitives

Primitive Answers
model_summary schema, units, storeys, entity types, property-set names
spatial_tree project → site → building → storey → space
elements_by_type elements of a type, subtypes included
elements_by_property property comparison; the fire-door workhorse
elements_missing_property data completeness: who has no value for this field
elements_missing_material no material through any of IFC's five ways of saying so
spaces_by_area rooms inside an area range, always in m²
property_values distinct values with counts — run this before guessing names
quantity_rollup totals grouped by type, storey or PredefinedType
element_detail expand specific GlobalIds to every pset and quantity

bimq queries prints their parameters. List queries return compact rows on purpose; element_detail is the drill-down, and keeping those separate is what stops a query from becoming the context dump it replaced.

Aggregates report their own coverage. A roll-up over 200 walls where 160 carry no quantity says so in summary and notes, because a total over 40 of 200 is not a total.

Policy

name = "consultant-readonly"

[allow_sources]
roots = ["/srv/bim"]
max_bytes = 536870912

[allow_queries]
queries = ["model_summary", "spatial_tree", "elements_by_type", "elements_by_property"]

[scope_storeys]
names = ["Level 2", "Level 3"]
include_unplaced = false

[allow_types]
types = ["IfcBuiltElement", "IfcSpace", "IfcBuildingStorey"]

[deny_properties]
properties = ["*Cost*", "Pset_Tender.*"]

[redact_properties]
properties = ["*.Owner*", "*SerialNumber*"]
placeholder = "[redacted]"

[max_results]
limit = 200
bimq policy check policy.toml   # validate before shipping
bimq rules                      # every rule, with an example

Notes on the design:

  • An unknown table is a hard error, not a warning. A file whose job is to withhold data must not fail open because of a typo.
  • deny and redact are different tools. A denied property is gone; a redacted one is present with a placeholder. The distinction matters to an agent: redaction says this exists and you are not being shown it, so the agent reports a gap instead of concluding nobody entered the data.
  • Denial covers the query side too. You cannot filter on a denied property, because op=gt value=1000 repeated a few times reconstructs it.
  • Withholding is reported, never silent. Answers carry policy.elements_withheld and a note. Truncation sets truncated: true.
  • Every answer is capped even with no policy at all. "Unlimited" is not a sane default for something feeding a context window.

Source formats

Format Notes
IFC-SPF (.ifc, .ifczip) IFC2X3 / IFC4 / IFC4X3, streaming reader, no dependencies
gbXML (.gbxml) energy models; ids are stamped gbXMLId, never confused with GlobalIds

Wanted, one per PR: Revit export (pyRevit/Dynamo JSON), Speckle stream, IFC-JSON, COBie. See CONTRIBUTING.md.

How the IFC reader stays small

bimq/sources/spf.py is a complete ISO 10303-21 reader in under 400 lines. The parts that matter:

  • The file is scanned in 4 MB blocks, so a 300 MB model is never one string. A block boundary can land inside a string literal, so the scanner explicitly matches unterminated literals and carries them forward. Tested at block sizes down to one byte, where the result must still be byte-identical.
  • A ; inside 'a;b' does not end a statement, '' is an escaped quote, and \X2\...\X0\ decodes to UTF-16 — so Phòng họp survives the round trip.
  • Comments appear between statements, inside parameter lists, and around section markers. All three are handled; the reported line still points at the entity.
  • Geometry is never loaded. An instance is kept only if its first attribute is a syntactically valid GlobalId — making it an IfcRoot subtype — or if it is one of ~30 unrooted carriers of property, quantity, material or unit data. The test is applied to the raw text before tokenising, which is where the parse time on a real file actually goes.

Check the throughput claim yourself without needing a model of your own — this writes a file shaped like a real export (a modest element count buried in geometry), parses it, and reports:

$ bimq bench --synthetic 20000
synthetic model: 20000 elements among 820006 instances
tmp6l05ix9x.ifc: 32.2 MiB, 20001 elements
parse: 1.71 s  ·  18.8 MiB/s  ·  11,680 elements/s
peak rss: 107 MiB  (3.3x file size)

820,006 instances go in; 20,001 elements stay resident. That ratio is the whole argument — resident size tracks how many things the building has, not how many points were needed to draw them.

Model files are treated as untrusted input. The gbXML reader refuses entity declarations outright, so a file cannot carry a billion-laughs expansion or an external entity pointing at /etc/passwd.

Units

Every number bimq returns is SI: metres, m², m³. A model authored in millimetres with areas in square metres (what Revit exports) and one authored in feet with areas in square feet both answer spaces_by_area max_m2=8 correctly. IfcConversionBasedUnit chains are resolved, not guessed.

Try it

The fixtures are synthetic — stated plainly, because a fixture pretending to be a real project is one nobody can check. What makes them useful is that the defects are deliberate and enumerated: a wall with no material, a fire door with no rating, a room below 8 m², a room with no area quantity at all, a door whose rating is inherited from its type, and a Vietnamese room name written with \X2\ escapes.

python scripts/make_fixture.py tests/fixtures

bimq describe tests/fixtures/office.ifc
bimq query elements_missing_material tests/fixtures/office.ifc type=IfcWall
bimq query spaces_by_area tests/fixtures/office.ifc max_m2=8
bimq query property_values tests/fixtures/office.ifc type=IfcDoor property=FireRating
bimq query spaces_by_area tests/fixtures/legacy-imperial.ifc max_m2=8   # authored in feet
bimq query spaces_by_area tests/fixtures/clinic.gbxml max_m2=8          # gbXML, same primitive

Development

make test        # the suite
make fixtures    # regenerate fixtures (byte-identical; CI checks this)
make bench       # parse throughput on a fixture

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选