migration-work

v6v7 · +11 −10 · View latest

⋯ 1200 unchanged lines
          <div class="case-node" aria-hidden="true"></div>
          <details class="case-disclosure">
            <summary class="case-summary" aria-controls="helpshift-hadoop-hive-to-spark-redshift-details">
              <span class="case-title" role="heading" aria-level="3">Helpshift<br>Hadoop + Hive → Spark + Redshift</span>
              <span class="case-preview">Supported 10x platform growth and improved data accuracy from 90% to 99.99% in one year.</span>
              <span class="case-meta">About 6 months · High complexity · Human-led delivery · Live data platform</span>
              <span class="case-title" role="heading" aria-level="3">Helpshift<br>Hadoop + Hive + HBase → Spark + Redshift</span>
              <span class="case-preview">Customer NPS rose from 4 to 25 because data became more accurate and refresh time fell from 24 hours to 1.5 hours.</span>
              <span class="case-meta">14 months from October 2021 · High complexity · Zero downtime · Human-led delivery</span>
            </summary>
            <div class="case-body case-content" id="helpshift-hadoop-hive-to-spark-redshift-details">
              <header class="case-header">
                <span class="case-status">Live data platform</span>
                <span class="case-status">Zero downtime</span>
              </header>
              <div class="case-metrics">
                <p class="metric" data-field="impact"><span class="metric-label">Business impact</span>Supported 10x platform growth, improved data accuracy from 90% to 99.99% in one year, and strengthened resiliency, scalability, and adoption.</p>
                <p class="metric" data-metric="duration"><span class="metric-label">Duration</span>About 6 months elapsed. The 1-year accuracy period is an outcome window, not modernization duration.</p>
                <p class="metric" data-metric="complexity"><span class="metric-label">Complexity</span>High. 2 billion devices, 1 PB daily ingestion, live accuracy requirements, and a broad warehouse change.</p>
                <p class="metric" data-field="impact"><span class="metric-label">Business impact</span>Customer NPS rose from 4 to 25 because data became more accurate and refreshed faster: accuracy improved from 90% to 99.99% in one year, and refresh time fell from 24 hours to 1.5 hours. The platform supported 10x growth while strengthening resiliency, scalability, and adoption. Helpshift's Chief Customer Success Officer tracked NPS.</p>
                <p class="metric" data-metric="duration"><span class="metric-label">Duration</span>14 months starting October 2021. The team had 4 people dedicated to the work: 2 senior and 2 junior. The one-year accuracy period is an outcome window, not modernization duration.</p>
                <p class="metric" data-metric="complexity"><span class="metric-label">Complexity</span>High. Old and new systems stayed live and continuously synced during a 16-pipeline transition spanning 2 billion devices and 1 PB daily ingestion. Customer retention and billing metrics could not pause or break; this was customer billing data, not Helpshift billing.</p>
              </div>
              <div class="case-facts">
                <section><h4>Why it was hard</h4><p>Moving Hadoop and Hive to Spark and Redshift had to preserve processing continuity and data correctness across 2 billion devices and 1 PB of daily ingestion.</p></section>
                <section><h4>How risk was controlled</h4><p>Stage processing and warehouse changes around Spark and Redshift, with Hudi, PostgreSQL, Arrow, Airflow, and DBT supporting continuity and correctness.</p></section>
                <section><h4>Why it was hard</h4><p>Hortonworks stopped supporting Hadoop, HBase, and Hive. Recurring HBase and Hive outages often forced teams to rebuild HBase data, while customer retention and billing metrics could not pause or break.</p></section>
                <section><h4>How risk was controlled</h4><p>Both systems stayed live and continuously synced. The team rewrote two critical pipelines on Spark and Redshift, supported by Hudi, PostgreSQL, Arrow, Airflow, and DBT, then ran old and new paths in parallel with data quality checks. Once the new outputs were trusted, the team shut down the old paths for those two pipelines and repeated the pattern for 14 more pipelines. The old system was then shut down with zero downtime.</p></section>
              </div>
              <p class="case-ai"><span>Delivery model</span>Human-led; no AI-assisted coding.</p>
              <p class="case-source"><span>Reference</span><a href="https://www.linkedin.com/in/akshay-vaidya-908702/" data-case-reference="akshay-vaidya" rel="noopener">Akshay Vaidya, Head of Engineering, Helpshift</a></p>
              <p class="case-judgment"><span>Human judgment</span>Senior architect judgment governed contracts, architecture, sequencing, gates, tradeoffs, and risk calls.</p>
              <p class="case-source"><span>Interview evidence</span>Kiran Kulkarni. <span>Reference</span><a href="https://www.linkedin.com/in/akshay-vaidya-908702/" data-case-reference="akshay-vaidya" rel="noopener">Akshay Vaidya, Head of Engineering, Helpshift</a></p>
            </div>
          </details>
        </li>
⋯ 287 unchanged lines