REDOOP 红象云腾
← Back to customer storiesOperations case / REDOOP CASE STUDY

Big data operations and maintenance solution

A customer runs multiple production and test clusters with thousands of nodes and close to 100 PB of capacity, requiring 7×24 operation and fast troubleshooting to locate and identify issues.

Big data operations and maintenance solution

Operations framework

Data processing flows are long and span many systems. When a business incident occurs, the related systems must be checked quickly, closing the gap between business and systems and continuously improving management and analysis capability.

Operations framework

Issue catalogue

Covers whole-cluster outages, node failures, slow performance, task and submission failures, abnormal data nodes, unbalanced allocation, missing data blocks, stalled tasks and slow access.

Outcome

Based on the customer’s actual needs and long-term maintenance, the cluster was optimized and improved in depth, producing a relatively complete solution for each class of issue.

This page is based on the publicly published case study on the original website; anonymous customers remain anonymous. Figures and results describe the original project.

LET’S BUILD WHAT’S NEXT

The next possibility starts here.

Talk to us about the Redoop big data platform that fits your business.

Talk to us