2018大数据技术指南(英文版)_38页-7mb
报告摘要
2018 DZone Guide to Big Data: Summary
Core Content
This document provides an overview of the Big Data landscape in 2018, highlighting the evolution of Big Data technologies, challenges, and future trends. It includes insights from DZone contributors and survey data on the adoption of tools, languages, and platforms in the Big Data domain.
Main Topics and Key Points
Introduction to Big Data
- Definition: Big Data is characterized by the three V's: Volume, Velocity, and Variety.
- Evolution: Initially seen as a fad, Big Data has matured and is now essential for modern business operations.
- Drivers: Technologies like Blockchain, Cloud, and IoT are enhancing the Big Data ecosystem.
- Challenges: Data contextualization, validity, noise, and abnormality are significant issues in handling Big Data.
Survey Highlights
- Participation: 540 software professionals participated in the survey.
- Demographics:
- 42% identify as developers or engineers.
- 23% are developer team leads.
- 54% have 10+ years of experience; 28% have 15+ years.
- 39% work in European-based companies; 30% in North American-based ones.
- Language Trends:
- Python has surpassed R as the dominant language for data science.
- 70% of respondents use Python for data science, while 50% use R.
- R is still popular for statistical analysis, but Python is catching up due to its ease of use and rich library ecosystem.
- Database Trends:
- MySQL remains the most popular DBMS, though its usage has slightly declined.
- Oracle and PostgreSQL have seen increased usage.
- NoSQL databases like MongoDB are widely used for Big Data, especially for handling semi-structured data.
Data Storage and Cloud Adoption
- Cloud Usage: 39% of respondents store data in the cloud, compared to 33% on-premise and 23% using hybrid solutions.
- Cloud Vendors:
- Amazon is the most popular (70%).
- Google Cloud and Microsoft Azure follow with 57% and 39%, respectively.
- Implications:
- Cloud adoption for Big Data is growing, but not as fast as in other areas like Continuous Delivery.
- The proximity of data can make handling "big" data more efficient.
Data Challenges
- Volume: 76% of respondents deal with large quantities of data.
- Velocity: 46% work with high-velocity data.
- Variety: 45% work with highly variable data.
- Common Challenges:
- Files and server logs are the most problematic data sources for volume and velocity.
- Relational and semi-structured data types are the most challenging.
- Data variety is mainly addressed through file-based data sources.
Blockchain and Big Data
- Overview:
- Blockchains are decentralized, peer-to-peer ledgers that securely store transaction data.
- They eliminate the need for middlemen and enable automated smart contracts.
- Benefits:
- Enhanced data security due to distributed architecture and encryption.
- Increased transparency and trust in data transactions.
- Potential for industry-wide data sharing and collaboration.
- Use Case:
- In agriculture, blockchain can track raw materials and products across the supply chain.
- It can help in verifying compliance and tracing issues like pests or fungi.
- Challenges:
- Setting up a blockchain network requires consensus on data sharing and format.
- It's a new model that may be counterintuitive for some industries.
HPCC Systems
- Overview:
- HPCC Systems is an open-source Big Data analytics platform.
- It is designed for efficiency and scalability, offering an end-to-end solution.
- Features:
- ETL Engine: Uses ECL, a powerful and easy-to-use programming language.
- Query and Search: Index-based search engine supporting SOAP, XML, REST, and SQL.
- Data Management Tools: Includes data profiling, cleansing, job scheduling, and automation.
- Predictive Modeling Tools: Supports linear and logistic regression, decision trees, and random forests.
- Strengths:
- One unified language (ECL) for data processing and analysis.
- Efficient handling of both structured and unstructured data.
- Proven performance and scalability for enterprise use.
Recommendations
- Data Handling:
- Use Apache Kafka for real-time data processing.
- Consider document store databases like MongoDB for handling semi-structured data.
- Cloud Strategy:
- Cloud storage is suitable for smaller enterprises with less upfront investment.
- Hybrid solutions can be a good compromise for sensitive data.
- On-premise solutions may be necessary for fast insights.
- Data Quality:
- Sanitize user inputs to maintain data integrity.
- Allocate sufficient time in project timelines for data preparation.
Conclusion
The guide aims to provide readers with insights into the current state of Big Data technologies and their applications. It emphasizes the importance of addressing data quality, leveraging cloud and open-source tools, and exploring the potential of Blockchain for enhancing data integrity and transparency. The document also highlights the growing adoption of Python in data science and the role of HPCC Systems in enabling efficient Big Data processing.
试读结束,高清完整版pdf/doc/ppt,请点下载