计算机行业智联汽车系列深度之39_小鹏vla2.0发布_智能驾驶体现更强大的泛化性_19页_1mb
报告摘要
Summary
Second-generation Vision-Language-Action (VLA) models, introduced by Xpeng, represent a significant advancement in multi-modal AI, moving beyond translation layers to directly map visual inputs to actions. Key developments include enhanced generalization in autonomous driving, such as handling small roads and recognizing human gestures. This model, based on Xpeng's Turing chip, supports low-bit width operations for efficient inference. The technology may spill over to other fields like robotics and low-altitude economy. Investment opportunities are associated with companies like Xpeng Auto, DeSer West, and others, but risks include competition from alternative approaches and valuation volatility.
展开完整摘要
试读结束,高清完整版pdf/doc/ppt,请点下载