Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
56f39d2
Mock up Host docs lifecycle IA
Jun 4, 2026
f2c0a90
Iterate CON-1518 host docs IA mockup
Jun 15, 2026
5840dde
Refine CON-1518 host docs mockup
Jun 17, 2026
c849fd7
Merge remote-tracking branch 'origin/main' into CON-1518-host-docs-ia…
Jun 17, 2026
ca24b9b
Add headless hosting guide
Jun 18, 2026
0cb28ff
docs(host): human-review common host questions
Jun 19, 2026
c649bcf
docs(host): fix hardware prep anchor jump
Jun 19, 2026
0525852
docs(host): surface common dashboard errors
Jun 19, 2026
d7d31ff
docs(host): add AER log search commands
Jun 19, 2026
45aadf6
docs(host): clarify host account signup path
Jun 19, 2026
8638b09
CON-1518 add GPU overview income estimate
Jun 22, 2026
c3bb70e
CON-1518 verify host fleet CLI docs
Jun 22, 2026
748ba51
CON-1518 clarify maintenance window commands
Jun 22, 2026
8a14c7f
CON-1518 fix anchored section offset
Jun 22, 2026
16db197
CON-1518 clarify XFS project quota setup
Jun 22, 2026
73861dd
CON-1518 clarify RAID monitoring expectations
Jun 22, 2026
71ead28
CON-1518 clarify canonical host setup flow
Jun 22, 2026
f373191
Add host machine error reference
Jun 24, 2026
64bd694
Clarify UDP machine error evidence
Jun 24, 2026
a8a935c
CON-1531 clarify vericode error meaning
Jun 24, 2026
a761af3
CON-1531 expand machine error guidance
Jun 24, 2026
f16f886
Stack headless host guide onto IA mockup
Jun 29, 2026
b81f5c2
Merge CON-1518 IA mockup into headless handoff
Jun 29, 2026
3c6ed79
Address host optimization guide coverage
Jun 29, 2026
818762b
Replace third-party host market chart links
Jun 29, 2026
a60a634
Document renter reports for hosts
Jun 29, 2026
05b1fa3
Clarify host diagnostics routing
Jun 29, 2026
1e702e0
Deduplicate host diagnostics reference
Jun 29, 2026
42d2968
Rename diagnostics page title
Jun 29, 2026
fcdd8ac
Update host diagnostics labels
Jun 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
62 changes: 52 additions & 10 deletions docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,7 @@
"group": "Teams",
"icon": "users",
"pages": [
"guides/teams/overview",
"guides/teams/teams-overview",
"guides/teams/managing-teams",
"guides/teams/teams-roles",
"guides/teams/legacy-teams"
Expand Down Expand Up @@ -588,26 +588,68 @@
"tab": "Host",
"groups": [
{
"group": "Concepts",
"icon": "lightbulb",
"group": "Before You Host",
"icon": "compass",
"pages": [
"host/hosting-overview",
"host/understanding-verification",
"host/earning"
"host/supported-hardware",
"host/persona-decision-guide",
"host/earning",
"host/guide-to-taxes"
]
},
{
"group": "Set Up",
"icon": "wrench",
"pages": [
"host/quickstart",
"host/account-hosting-agreement",
"host/hardware-prep",
"host/storage-setup",
"host/network-ports",
"host/installing-host-software",
"host/headless-install",
"host/vms"
]
},
{
"group": "Guides",
"icon": "warehouse",
"group": "Verify & List",
"icon": "badge-check",
"pages": [
"host/optimization-guide",
"host/verification-stages",
"host/understanding-verification",
"host/how-to-self-test",
"host/market-metrics",
"host/self-test-reference",
"host/pricing-your-listing",
"host/not-in-search"
]
},
{
"group": "Operate",
"icon": "gauge",
"pages": [
"host/first-24-hours",
"host/reliability-uptime",
"host/maintenance-windows",
"host/removing-recreating-machines",
"host/fleet-operations",
"host/common-errors-diagnostics",
"host/machine-errors",
"host/datacenter-status",
"host/payment",
"host/guide-to-taxes",
"host/vms"
"host/workload-policy"
]
},
{
"group": "Reference",
"icon": "book-open",
"pages": [
"host/common-host-questions",
"host/glossary",
"host/hosting-agreement",
"host/community"
]
},
{
Expand Down Expand Up @@ -808,7 +850,7 @@
"links": [
{
"label": "FAQ",
"href": "/guides/reference/faq"
"href": "https://docs.vast.ai/guides/reference/faq"
},
{
"label": "Discord",
Expand Down
50 changes: 50 additions & 0 deletions host/account-hosting-agreement.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
---
title: "Account & Hosting Agreement"
sidebarTitle: "Account & Agreement"
description: "Host account conversion and agreement flow."
"canonical": "/host/account-hosting-agreement"
personas:
- pro-operator
- headless-operator
- business-owner
- hobbyist
---

<div className="persona-chips"><span className="persona-chip">All host personas</span></div>

Use this page when a host account, hosting agreement, or Machines tab is not behaving as expected.

<a id="separate-host-account" />
## Do I need a separate host account?

Yes. Use a dedicated account for hosting. Do not use the same account for both client rentals and host operations, because that setup is unsupported and can cause account and machine-management issues.

## How to accept the hosting agreement

Once your host account is created, open the [host setup page](https://cloud.vast.ai/host/setup/). There is a link in the first paragraph to the hosting agreement. Read through the agreement. Once you accept, your account is converted to a hosting account, and a Machines link appears in the navigation. Your account can now list machines that are running the daemon software.

The setup page is the canonical place for the current agreement link and generated install command. Do not rely on copied buttons, screenshots, or old setup instructions.

<a id="host-features-tab" />
## What must happen before I can see host features or the Machines tab?

Your account must be enabled for hosting and the hosting agreement must be accepted. After the account is converted to a host account, the Machines navigation and other host features should become available.

If the account still behaves like a client account, confirm that you are signed into the intended host account and that the agreement flow completed.

<a id="agreement-stuck" />
## What if I accepted the agreement but still see a client account?

Confirm that you are signed into the correct dedicated host account, then refresh or sign out and back in. If the Machines tab is still missing, treat it as an account-state issue rather than a Linux install issue. Contact Vast support with the host account email and screenshots or error messages from the setup flow.

## Stale or deprecated setup links

Copy the install command only from the official setup page while signed into the host account. Commands from old notes, screenshots, third-party guides, or another account can contain expired or account-specific material. See [Installing Host Software](/host/installing-host-software#install-command).

## Related pages

| Topic | Read next |
| --- | --- |
| First-time setup path | [Hosting Quickstart](/host/quickstart) |
| Installing the host software | [Installing Host Software](/host/installing-host-software) |
| Agreement reference | [Hosting Agreement](/host/hosting-agreement) |
207 changes: 207 additions & 0 deletions host/common-errors-diagnostics.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,207 @@
---
title: "Host Diagnostics"
sidebarTitle: "Diagnostics"
description: "Host log collection and diagnostic evidence paths."
"canonical": "/host/common-errors-diagnostics"
personas:
- pro-operator
- headless-operator
---

<div className="persona-chips"><span className="persona-chip">Pro Operator</span><span className="persona-chip">Headless / DC</span></div>

Use this page to collect logs and host-side evidence. For exact Machines-page error strings and what they usually mean, start with [Machine Error Reference](/host/machine-errors).

## Where To Start

| Situation | Start here |
| --- | --- |
| Machines page shows a red error, `vericode=8`, or a host-facing runtime message | [Machine Error Reference](/host/machine-errors) |
| Self-test failed and you need a support bundle | [Logs and support bundles](#logs) |
| You need host daemon, service, GPU, kernel, storage, or port evidence | Use the diagnostic commands on this page |
| Port reachability is the main issue | [Network & Ports](/host/network-ports) |
| Renter reported a machine issue | [Workload Policy: Renter Reports And Host Logs](/host/workload-policy#renter-reports-and-host-logs) |

<a id="logs" />
## Logs And Support Bundles

Installer logs are written to `vast_host_install.log` in the directory where you launched the installer, not under `/var/lib/vastai_kaalia`.

```bash
cat vast_host_install.log
```

If the installer created a compressed log archive:

```bash
tar -xzvf vastai_install_logs.tar.gz
cat vast_host_install.log
```

The host daemon log is:

```bash
sudo tail -n 100 /var/lib/vastai_kaalia/kaalia.log
```

For self-test failures, the CLI can create a diagnostic bundle. The normal command is:

```bash
vastai self-test machine <machine_id>
```

Failure bundles are saved by default under:

```text
/tmp/vast_selftest_<machine_id>_<timestamp>.tar.gz
```

You can override the output directory:

```bash
vastai self-test machine <machine_id> \
--support-bundle-dir /path/to/output
```

Bundles can include:

- `self-test-output.log`
- `self-test-result.json`
- `manifest.json`
- `collection-errors.json`
- `instance/show-instance.json`
- `instance/container.log`
- `instance/daemon.log`

<a id="host-service-snapshot" />
## Host Service Snapshot

For a quick SSH check, collect service state, recent service logs, daemon logs, and the configured host port range:

```bash
systemctl is-active vastai.service vast_metrics.service docker nvidia-persistenced.service
sudo journalctl -u vastai.service -n 80 --no-pager
sudo journalctl -u vast_metrics.service -n 80 --no-pager
sudo tail -n 100 /var/lib/vastai_kaalia/kaalia.log
sudo cat /var/lib/vastai_kaalia/host_port_range
```

Services should be `active`, logs should not show restart loops or repeated fatal errors, and the configured port range should match the forwarded ports.

<a id="gpu-kernel-logs" />
<a id="gpu-pcie-issue" />
<a id="nvidia-smi-fails" />
<a id="nvml-mismatch" />
<a id="gpu-falls-off-bus" />
<a id="bad-bandwidthtest2" />
<a id="failed-cdi" />
## GPU And Kernel Diagnostics

Use these commands to collect GPU, PCIe, AER, Xid, driver, and Docker GPU-injection evidence after an exact error has pointed you toward GPU or runtime health. For error meanings and first-response guidance, use [Machine Error Reference](/host/machine-errors).

- **NVRM**: NVIDIA kernel driver messages.
- **Xid**: NVIDIA GPU fault, reset, or error codes.
- **PCIe**: the bus/link between the GPU, motherboard, and CPU.
- **AER**: PCIe Advanced Error Reporting messages. Repeated AER messages can point to risers, slots, power, BIOS lane settings, cabling, motherboard, or GPU hardware issues.

Check the current boot:

```bash
sudo journalctl -k -b --no-pager | grep -Ei 'NVRM|Xid|AER|PCIe|fallen|GPU has fallen'
sudo dmesg -T | grep -Ei 'NVRM|Xid|AER|PCIe|fallen|GPU has fallen'
```

Search kernel and syslog history for AER/PCIe evidence:

```bash
sudo journalctl -k -b --no-pager | grep -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid'
sudo journalctl -k -b -1 --no-pager | grep -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid'
sudo dmesg -T | grep -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid'
sudo grep -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid' /var/log/syslog /var/log/syslog.1 /var/log/kern.log /var/log/kern.log.1 2>/dev/null
sudo zgrep -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid' /var/log/syslog.*.gz /var/log/kern.log.*.gz 2>/dev/null
```

While reproducing under load, watch live kernel messages in another SSH session:

```bash
sudo journalctl -kf | grep --line-buffered -Ei 'AER|PCIe Bus Error|pcieport|NVRM|Xid'
```

Confirm host GPU visibility and Docker GPU injection:

```bash
nvidia-smi -L
sudo docker run --rm --gpus all nvidia/cuda:12.2.0-base-ubuntu22.04 nvidia-smi -L
```

If the machine is idle and not rented, you can also run a short load test with a GPU burn image you trust:

```bash
sudo docker run --rm --gpus all oguzpastirmaci/gpu-burn 60
```

Correlate timestamps with self-test runs, rentals, `gpu-burn`, or other load tests. Repeated messages during load are much more concerning than a small number of corrected messages during boot.

<a id="red-error" />
<a id="port-issue" />
<a id="docker-cache" />
## Machine Error Lookup

Use the Machine Error Reference for exact message meanings, likely causes, and next checks:

| Message or symptom | Canonical reference |
| --- | --- |
| Red machine error, unhealthy, or de-verified machine | [Generic red machine error](/host/machine-errors#generic-red-machine-error) |
| `Port issue` or `Port Networking Issues` | [Port Networking Issues](/host/machine-errors#port-networking-issues) |
| TCP works but UDP fails | [TCP works but UDP fails](/host/machine-errors#tcp-works-but-udp-fails) |
| `GPU PCIE issue`, AER, PCIe, Xid 79, or fallen off bus | [GPU PCIe Issue](/host/machine-errors#gpu-pcie-issue) |
| `failed to inject CDI devices` or `unresolvable CDI devices` | [CDI Device Injection](/host/machine-errors#cdi-device-injection) |
| `bad bandwidthtest2` | [`bad bandwidthtest2`](/host/machine-errors#bad-bandwidthtest2) |
| `nvidia-smi` fails, NVML mismatch, or GPU disappears | [GPU Missing Or Unhealthy](/host/machine-errors#gpu-missing-or-unhealthy) |
| Full client storage, `no space left on device`, or missing expected disk space | [Full Client Storage](/host/machine-errors#full-client-storage) |

## Storage And Port Evidence

For storage or port reports, collect the host-side facts before changing machine state:

```bash
sudo cat /var/lib/vastai_kaalia/host_port_range
df -h /var/lib/docker
sudo docker system df
findmnt /var/lib/docker -no SOURCE,FSTYPE,OPTIONS
```

For networking cases, also capture the configured router, firewall, datacenter ACL, or provider forwarding state. For the full network checklist, see [Network & Ports](/host/network-ports). For storage meanings and cleanup guidance, see [Machine Error Reference: Full Client Storage](/host/machine-errors#full-client-storage).

<a id="collect-logs" />
## Before Asking For Help

Include:

- Account context and exact command.
- The affected machine record in the console, relevant offer/contract context if support asks, and whether the account is the host account.
- Exact timestamps with timezone.
- Exact error strings, screenshots, or CLI output.
- What changed recently: reboot, driver, kernel, Docker, BIOS, storage, router, or ISP.
- Host OS, GPU model/count, NVIDIA driver version, and Docker storage mount output.
- Tested external IP and port for networking issues.
- Self-test support bundle when available.
- Installer log, daemon log, and relevant kernel log excerpts.

Review bundles before sharing them. Share sensitive machine/account details only in the appropriate support channel.

<a id="escalate-support" />
## Escalate To Vast Support

Escalate account conversion, hosting agreement, payout, API permission, backend machine record, ghost machine, suspected backend bug, or platform-state mismatch issues.

Local Ubuntu, Docker, GPU driver, hardware, power, thermals, and consumer networking are primarily host responsibilities. Vast hosting requires direct public inbound TCP/UDP reachability; CGNAT or double NAT without a real public forwarding path is not a supported hosting network setup.

## Related Pages

| Topic | Read next |
| --- | --- |
| Install failures | [Installing Host Software](/host/installing-host-software#install-failures) |
| Machine error strings | [Machine Error Reference](/host/machine-errors) |
| Self-test errors | [Self-Test Reference](/host/self-test-reference) |
| Connectivity | [Network & Ports](/host/network-ports) |
Loading