Senior Data R&D Engineer

  • Hong Kong, Hong Kong
  • Full-Time
  • On-Site

Job Description:

Job Overview

Responsible for the construction of financial and Web3 data infrastructure tailored for AI application scenarios—building a complete end-to-end chain from external data ingestion and standardized processing to metrics analysis and data services. Capable of independently taking ownership of key modules (technical architecture design to final delivery) and continuously driving improvements in data quality, system stability, and engineering efficiency.

Responsibilities

  • Data Architecture & Platform Engineering: Lead core module design and evolution focused on multi-source heterogeneous data ingestion, time-series data processing, and AI data consumption scenarios. Establish scalable data models, ingestion standards, and service contracts while balancing delivery speed, maintenance costs, and long-term scalability.
  • Multi-Source Data Ingestion & Governance: Direct the integration of stock/ETF market feeds, macroeconomic stats, financial news, social sentiment, on-chain data, and DeFi/derivatives metrics. Build reusable collection, parsing, and validation capabilities to resolve multi-source asset mapping, methodology discrepancies, missing data, and historical revisions.
  • Data Processing & Analytics: Translate investment research requirements into explicit data models, metric formulas, and computation logic. Build aggregated analytics and event identification for price anomalies, capital flows, market sentiment, and key market events to empower signal generation, explainability, and backtesting.
  • AI Data Service Layer: Collaborate with product, investment research, and AI/algorithm teams to design data interfaces suited for AI retrieval and tool calling (Function Calling). Refine management mechanisms for raw facts, derived metrics, and AI analysis outputs to guarantee low latency, lineage attribution, and analytical traceability.
  • Performance & Reliability: Manage collection scheduling, concurrency control, idempotent writes, crash recovery, historical backfilling, and data replay mechanisms. Establish measurable data quality/service metrics (SLAs) while constantly optimizing throughput, query latency, storage overhead, and external API call costs.
  • Technical Delivery & Engineering Practices: Own the end-to-end delivery lifecycle—from requirement refinement and architectural design reviews to dev/testing, deployment, and ongoing maintenance. Enhance overall team delivery quality and speed via rigorous code reviews, technical documentation, and AI-assisted software engineering practices.

Qualifications

  • Core Experience: 3+ years of Java backend development experience with a track record of independently leading core modules or complex projects. Proven ability to identify core problems, define system boundaries, evaluate design trade-offs, and drive execution to successful delivery even under ambiguous requirements.
  • Java & System Design Mastery: Deep understanding of concurrency models, thread pools, JVM memory/GC tuning, transaction management, and exception handling. Proficient with Spring Boot; capable of designing modular, easy-to-test, and evolvable services using Java 21 and Spring Boot 3, alongside diagnosing complex performance and stability bottlenecks.
  • Data Modeling & Database Optimization: Proficient in PostgreSQL and JPA/Hibernate with a solid grasp of indexing, transaction isolation levels, locking mechanisms, and query execution plans. Experienced in designing time-series partitioning, batch writes, Upserts, and aggregation queries based on scale and access patterns to resolve performance and consistency issues caused by rapid data growth, concurrent writes, and historical backfills.
  • Data Ingestion & Governance: Experience designing multi-source data adapter and standardization frameworks that systematically handle API schema changes, rate limits/quotas, incremental cursors, duplicate/out-of-order data, and job fault tolerance. Deep understanding of asset tickers, numeric precision, temporal alignment, data lineage, and data quality validation to build verifiable and replayable data pipelines.
  • Distributed Systems Engineering: Proficient with Redis, Dubbo, Nacos, SchedulerX, or equivalent distributed infrastructure. Sound judgment on the operational boundaries of cache consistency, distributed locking, timeout retries, and fault isolation. Skilled in designing task sharding, concurrency control, observability/alerting, and capacity planning around business goals.
  • Data Analysis & Business Abstraction: Hands-on experience in traditional finance, Web3, or both. Deep understanding of the business domain behind financial metrics, underlying data sources, and practical constraints. Proficient in using SQL, scripts, and sample verification to validate analytical conclusions and identify metric drifts, point-in-time mismatches, and data anomalies affecting research signals.
  • AI-Assisted Engineering: Hands-on experience using AI coding tools (e.g., Codex, Claude Code) in daily development. Skilled at organizing prompt requirements, architectural constraints, and code contexts to perform task decomposition, code generation, unit test creation, debugging, and code reviews. Capable of catching logical bugs, edge-case omissions, and design flaws in generated code, verifying quality through testing, and taking final responsibility for delivered features.
  • Collaboration & Leadership: Strong technical communication skills to clearly articulate designs and risks. Adept at cross-functional collaboration with product, algorithm, and upstream/downstream engineering teams to resolve complex integration issues. Passionate about automated testing, observability, and post-mortems, with a habit of abstracting project lessons into reusable tools, frameworks, and engineering standards.

Preferred / Bonus Points

  • Experience building market data feeds or investment research platforms, with deep insights into trading calendars, stock split/dividend adjustments, historical revisions, point-in-time (PIT) data, and look-ahead bias prevention.
  • Experience with on-chain data indexing, address behavior profiling, transaction flow tracking, or DeFi/derivatives data analytics—capable of handling multi-chain asset mapping, block confirmation strategies, and chain reorganizations (reorgs).
  • Hands-on delivery experience with AI Agents, RAG pipelines, Tool Calling, or signal analysis products, including designing data attribution, evaluation metrics, and quality feedback loops.
  • Experience establishing reusable AI development workflows—such as building custom project prompts/instructions, custom Skills, Model Context Protocol (MCP) integrations, or integrating AI coding tools into CI/CD, testing, and code review processes with measurable productivity gains.
  • Experience taking point on technical architecture, driving cross-team engineering initiatives, or mentoring junior-to-mid-level engineers.