FUNGIBLE-带你了解以数据为中心的DPU处理器(英文)-2020.11-25页_1mb
报告摘要
Fungible DPU™: A New Category of Microprocessor for the Data-Centric Era
Core Mission
Revolutionize the performance, economics, reliability, and security of all scale-out data centers through the introduction of a new category of microprocessor: the Fungible DPU™.
Context and Challenges
Modern data centers face several critical challenges:
- Large footprint (power and space): Inability to pool expensive resources and inefficient execution of data-centric computations.
- Scaling challenges: Handling both very small and very large systems.
- Increasing complexity: Technology limits and growing security vulnerabilities.
The Fungible DPU™ is designed to address these challenges by focusing on data-centricity as the core driver of architecture.
Key Problems in Data Centers
The five root causes of inefficiency in data centers are:
- Inefficient data interchange between nodes
- Unreliability
- Inefficient execution of data-centric computations inside nodes
- Per-context state
- Inflexibility
These issues are exacerbated by the fact that data-centric applications:
- Arrive as packets
- Require frequent context switching
- Involve modification of state
- Have I/O-dominated arithmetic and logic
- Are increasingly insecure
Fungible DPU™ Addresses All Five Root Causes
The Fungible DPU™ is designed to solve the above problems by:
- Optimizing data movement and processing
- Providing secure execution environments
- Enabling efficient data-centric computations
- Supporting flexible and scalable state management
- Offering programmable and adaptable architectures
Architecture Overview
The Fungible DPU™ is composed of two main clusters:
8 Data Clusters
- 192 processor threads
- Full cache coherency
- Tightly integrated accelerators
- 6 cores * 4 threads
- Multi-threaded accelerators for:
- Data movement
- Data lookup
- Data security
- Data reduction
- Data protection
- Data analytics
Control Cluster
- 8 processor threads
- Runs the control plane on Linux
- Includes:
- Secure Enclave: Provides secure boot, key vault, and binary signing
- Public Key Crypto Engines: Supports RSA and elliptic curve cryptography
- True Random Number Generator (TRNG)
- Physically Unclonable Function (PUF)
- Cluster Cache & Memory Manager
Key Components
- High Bandwidth HBM2 Memory: 8GB, 4Tbits/sec, integrated in the package
- High Capacity DDR4 Memory: 2x DDR4 controllers, ECC enabled, up to 2666 MHz, 512GB capacity
- Flexible Network Engine (800G): Implements TrueFabric™ endpoint, supports low latency Ethernet, L2/L3/L4 forwarding, and end-to-end encryption
- Flexible Host Engine (512G): Includes 16 independent dual-mode controllers, supports SR-IOV, and provides hardware virtualization
Data Path Programming Model
- MIPS-64 Hardware Threads execute run-to-completion C-code
- Heterogeneous Accelerator Threads handle specialized data processing tasks
- On-chip fabric enables tight coupling between data and control planes
Software Programmability
- Multiple levels of programmability: Includes DPU control and data planes, eBPF for host-side data path code execution, and Northbound APIs for orchestration systems
- Fully general programmability: All code is written in ANSI-C
- Fast thread switching and tight integration with accelerators ensure no performance compromises
Performance Metrics
| Service | Measured Performance | Estimated Performance |
|---|---|---|
| TCP (Single Flow, Multi Flow) | 50Gbps, 250Gbps | 70Gbps, 400Gbps |
| TLS Session Setup Rate | 32,000/sec | 100,000/sec |
| IPSEC (Single Flow, Multi Flow) | - | 10Gbps, 250Gbps |
| Stateful Firewall | - | 370Gbps |
| OVS | - | 400Gbps |
| Load Balancer | 256Gbps | 300Gbps |
| Block Store (4K IOPS) | 8M | 10M |
| Video Streaming | 256Gbps | 300Gbps |
| TPC-H Benchmark (relative to X86) | 3X–100X | - |
TrueFabric™ Performance
| Traffic Pattern | Fabric Utilization | Latency Mean | Latency Variance | Latency P99 |
|---|---|---|---|---|
| 1024 * (Node to Node) | 90.7% | 1.84μs | 0.13μs | 2.14μs |
| 1024 Node to 1024 Node | 93% | 2.10μs | 0.32μs | 3.30μs |
| 1024 Nodes to 1 Node | 90% | 1.71μs | 0.12μs | 1.75μs |
Network Configuration
- 1024 nodes with 200Gbps/Node
- Two-tier leaf-spine architecture
- Leaf ZLL: 500ns
- Spine ZLL: 500ns
- iMix packet profile for performance testing
Fungible DPU™ as a New Microprocessor Category
The Fungible DPU™ is purpose-built for the data-centric era, offering:
- Data-centric architecture: Multi-core, MIMD + tightly-coupled accelerators
- High throughput for multiplexed workloads
- TrueFabric™ enables disaggregation and pooling of resources
- Specialized memory system and on-chip fabric for optimized performance
- Ideal for network, storage, security, and virtualization applications
- Data-centric computations run >10X more efficiently compared to traditional architectures
Product Launch
Fungible announced the F1 and S1 DPUs, both sharing a common architecture and programming model. These DPUs are designed for:
- Storage targets
- AI servers
- Security appliances
- Analytics applications
They support:
- Bare-metal virtualization
- Storage initiator and local instance storage
- NFV applications
- Node security
Conclusion
The Fungible DPU™ represents a new paradigm in microprocessor design, specifically tailored for the data-centric era. By addressing the root causes of inefficiency in data centers, it offers enhanced performance, security, and flexibility, making it an ideal solution for network, storage, security, and virtualization workloads.
试读结束,高清完整版pdf/doc/ppt,请点下载