What You Actually Need to Know Before Touching ONefs
The ONefs Admin Guide is a dense, somewhat frustrating document from Huawei that covers the fabric orchestration layer for their data center switching architecture. It assumes you already understand their switch platforms, CLI conventions, and the general concept of network fabric orchestration. If you're starting from zero, you'll spend more time cross-referencing than actually configuring anything. I recommend keeping a second browser tab open with the hardware installation guide so you're not jumping between PDFs constantly. I spent about three weeks last year doing a greenfield ONefs deployment for a mid-size data center. The guide itself is roughly 400 pages across multiple chapters, and quite a few sections read like they were translated from Chinese without a networking-specific glossary. You'll run into awkward phrasing, incomplete command examples, and the occasional step that references a feature flag that doesn't exist yet in your firmware version.
Onefs Admin Guide — Where to Get It and What It Actually Covers
You can download the current release directly from the Huawei support portal at support.huawei.com. You'll need an account and your device serial numbers verified before the download links become visible. The guide gets updated quarterly alongside software releases, so make sure the version number in the footer matches your ONefs OS build. Mismatched documentation is the most common source of "this command doesn't work" tickets I see in the forums. The guide is organized into setup, configuration, validation, and troubleshooting sections. The setup chapter walks through initial appliance deployment and license activation. The configuration section covers fabric creation, VLAN assignment, overlay network definition, and integration with existing IP infrastructure. The troubleshooting chapter is where most people end up, because the earlier sections tend to describe ideal-state workflows rather than what happens when something goes wrong during provisioning.
Configuring a Fabric — The Part the Guide Gets Mostly Right
When you're defining your fabric topology, the guide provides reasonable step-by-step instructions. You add your leaf and spine switches to the orchestration plane, assign them roles, and then the system auto-discovers the physical topology. This works well in clean lab environments. In production, it tends to misidentify link speeds on certain QSFP28-to-Copper DAC combinations, which causes the fabric to enter a degraded state until you manually override the port parameters. The workaround I use is straightforward. Before declaring the fabric ready, I pull the actual port speed and autonegotiation status from each switch using the CLI. I compare it against what ONefs has recorded. Any mismatch gets corrected manually in the inventory view. It adds about 20 minutes to a standard deployment but prevents a class of issues where traffic paths look healthy in the controller but actually blackhole under load. I should mention that the fabric creation wizard has a known behavior where it assigns spine-to-leaf interconnects to the default forwarding instance even when you've configured custom VXLAN segments. The guide mentions this as a footnote but doesn't emphasize it enough. If you're running multi-tenant overlays, always verify the forwarding instance assignment after the initial fabric build completes. Otherwise you'll spend hours debugging tenant isolation issues that trace back to a default-VNI leak.
Get the Full Details

License Management — Where Beginners Usually Stumble
ONefs licensing is modular and tied to node count, feature set, and bandwidth tiers. The Admin Guide explains this but doesn't make clear that certain features are gated behind premium licenses even when the software appears to accept the configuration command. You'll configure a policy, get no errors, and then watch it silently fail to apply at runtime. The license state screen shows exactly which features are active, but it's buried in the management UI under a non-obvious menu path. Here's the practical fix: before provisioning anything beyond basic Layer 2, run a license inventory check and capture the output. This gives you a baseline. After each major configuration change, run it again. If a feature stops working and the license was valid before, you know the new configuration hit an unlicensed path. This saved me from what would have been a very expensive support call.
Troubleshooting and Validation
The troubleshooting section of the guide covers the standard diagnostics — link status verification, topology reconciliation, and service health checks. What it doesn't cover well is the interaction between ONefs and external BGP route reflectors, which is a common deployment pattern. When your route reflectors reset or flap, ONefs can enter a state where it retains stale EVPN routes for 15 to 20 minutes before reconvergence. During that window, traffic forwards along suboptimal or broken paths. The configuration parameter that controls this behavior is the BGP hold-time and keepalive interval on the underlay interfaces. The guide recommends default values, but in production environments those defaults cause exactly this kind of delayed reconvergence. Adjusting the hold-time to 9 seconds with a 3-second keepalive interval cuts the reconvergence window to roughly 10 seconds in my experience. This is documented in the performance tuning appendix but easily missed if you're reading the guide cover to cover linearly. Another thing the guide handles inadequately: firmware rollback procedures. If an ONefs upgrade introduces a regression in your particular switch model — and this happens more often than Huawei would publicly acknowledge — the rollback process requires a specific sequence of CLI commands that isn't clearly laid out in the main admin guide. You'll find it in a separate advisory document linked from the release notes page. Plan for this before you push an update to production.
Common Configuration Mistakes
I've seen three recurring mistakes that anyone following this guide should avoid. First, applying overlay configurations before the underlay fabric has fully converged. The system allows it, and you'll get a clean success message, but the VXLAN tunnels won't carry traffic correctly until the underlay stabilizes. Wait at least 10 minutes after fabric creation before provisioning overlays. Second, mixing switch models from different hardware generations within the same fabric without verifying feature parity. Some LACP and QoS features behave differently across generations, and the guide treats them as identical. Third, skipping the post-deployment validation script. The built-in health check catches about 60 percent of configuration errors. Running the optional comprehensive validation catches most of the rest. The guide is serviceable if you approach it as a reference rather than a tutorial. It works best when you have a working deployment and need to look up a specific procedure. Reading it cover to cover as your primary learning resource will slow you down considerably. Pair it with hands-on lab time and you'll pick up the operational reality faster than the documentation will teach it to you.
