Maksym Diachuk is a Senior Software Engineer, AI Infrastructure researcher, and software architect with more than nine years of professional experience designing large-scale distributed systems, AI infrastructure, and high-performance data processing platforms. His work focuses on building scalable AI-driven architectures that transform historical documents into structured, searchable, and knowledge-rich digital collections through machine learning, optical character recognition (OCR), natural language processing (NLP), large language models (LLMs), cloud-native technologies, and distributed computing.
Since joining MyHeritage in 2019, Maksym has contributed to the design and implementation of AI-powered systems supporting one of the world’s largest historical genealogy platforms. During this period, the historical records collection expanded from approximately 10 billion to more than 32 billion records. He designed and deployed production-grade AI pipelines that extracted approximately 11.6 billion searchable records from historical OCR documents and developed automated post-processing systems that generated more than 1.3 billion searchable records from over 25,000 U.S. City Directory volumes while maintaining privacy-preserving data processing standards.
Maksym also contributed to the AI infrastructure supporting the publication of the 1950 United States Census, helping design scalable processing pipelines that enabled approximately 150 million census records to be prepared and published within approximately 13 hours after the official release. In addition, he contributed to the backend infrastructure powering OldNews, enabling full-text search across more than 100 million historical newspaper pages while supporting continuous large-scale content ingestion.
His professional engineering experience and subsequent systems-level research have contributed to the development of an original methodology for large-scale AI-driven digitization of historical records. This methodology integrates modern advances in OCR, NLP, large language models, Data Engineering, cloud infrastructure, distributed computing, automated quality assurance, semantic enrichment, knowledge representation, and provenance into a unified architectural approach for transforming heterogeneous archival collections into structured, searchable, and interoperable knowledge systems.
Maksym is the author of the technical monograph “
AI-Powered Digitization of Historical Records: Architectures, Pipelines, and Scaling Strategies” (Gutenberg Publishing House, 2026; ISBN 978-617-8660-48-2; DOI 10.66481/978-617-8660-48-2), in which he presents an original systems-level methodology for scalable AI-driven historical data transformation. The monograph draws on his experience designing production systems for large-scale historical data processing and presents architectural principles, processing pipelines, and scaling strategies for integrating OCR, NLP, large language models, cloud computing, automated quality assurance, metadata enrichment, and distributed processing into production-grade AI infrastructure.
The core concepts of this methodology are also reflected in Maksym’s scholarly and professional publications dedicated to AI Infrastructure, large-scale historical data processing, and the digital transformation of historical archives.
His research interests include AI Infrastructure, Software Architecture, Distributed Systems, Cloud Computing, Data Engineering, Knowledge Graphs, Generative AI, intelligent document processing, knowledge representation, digital preservation, and scalable information systems.