RESEARCH AGENDA / SCIENTIFIC COMPUTING

Numerical foundations. Modern systems.

Bridging numerical analysis, software architecture, and high-performance computing to build verifiable, scalable simulation software for exascale systems.

01 /Scientific solver modernization

Read full dossier
ACTIVE MODERNIZATIONCODEBASE: IZEM CFD
TARGET: TOUBKAL HPCTRACK 01

Track 01 · Scientific Computing & Solver Architecture

Modernization of a Reactive Compressible Multi-Species Solver

Decomposing legacy Fortran into a modular C++17 architecture with Cantera thermodynamics, WENO-5/TENO shock-capturing, and parallel HDF5 scientific I/O.

L1Physics: Cantera Kinetics
L2Numerics: WENO-5 / TENO
L3Systems: C++17 / HDF5
L4HPC: Kokkos / MPI
Explore Research Dossier & Architecture
  1. 01

    Understand & Audit

    Legacy Fortran Solver

    Baseline verified
  2. 02

    Decouple & Interface

    Architecture Redesign

    Interfaces defined
  3. 03

    Rebuild & Modernize

    Modular C++17 / Julia Core

    Active development
  4. 04

    Verify & Validate

    Verification & Validation

    Continuous testing
  5. 05

    Scale & Port

    HPC Scale & Portability

    HPC target
  • Fortran
  • C++17
  • Cantera
  • WENO-5
  • TENO
  • HDF5
  • MPI
  • Kokkos
VERIFIED AGAINST LEGACY REFERENCE

02 /Exascale observability & heterogeneous systems

RESEARCH MAP / TRACK 02

Tuning, debugging, and monitoring heterogeneous systems at exascale.

Track 2 focuses on the tools and methods needed to understand heterogeneous HPC systems where CPUs coordinate thousands of nodes and GPUs provide much of the computational power.

CPU + GPU×MPI / OPENMP / CUDA×OBSERVABILITY
TRACK 02 / HETEROGENEOUS HPCRESEARCH WORKFLOW
TRACK 02HETEROGENEOUS HPCPERFORMANCE / BEHAVIORCPU + GPUHARDWAREMPI / CUDAMODELSTRACINGLOW OVERHEADDEBUGGINGROCGDBVISUALIZESCALE

03 /Research interests

Investigating how to structure, coordinate, and accelerate numerical software across modern parallel systems—from cache line vectorization to exascale cluster execution.

01SYSTEMS AT SCALE

High-Performance Computing

Supercomputing Infrastructure

Designing simulation software to effectively exploit massively parallel supercomputing resources, batch schedulers, and high-speed interconnects.

  • MPI
  • SLURM
  • Toubkal HPC
  • InfiniBand
02CONCURRENCY

Parallel Computing

Multi-Core & Decomposition

Decomposing spatial computational domains and coordinating multi-threaded execution with minimal synchronization friction.

  • OpenMP
  • Pthreads
  • Domain Decomposition
03RUNTIME DAGS

Task-Based Parallelism

Dynamic Scheduling

Expressing scientific workflows as Directed Acyclic Graphs (DAGs) to enable latency hiding, work-stealing, and asynchronous execution.

  • Task DAGs
  • HPX
  • Dependency Tracking
04ACCELERATORS

Heterogeneous Computing

CPU + GPU Co-Design

Mapping scientific numerical kernels across host CPUs and throughput-oriented GPU accelerators, optimizing data movement across PCIe and NVLink.

  • CUDA
  • ROCm
  • Unified Memory
  • PCIe Transfer
05COMMUNICATION

Distributed Computing

Inter-Node Scalability

Structuring low-overhead ghost-cell halo exchange, non-blocking collective communication, and data locality across distributed cluster nodes.

  • Non-blocking MPI
  • Halo Buffers
  • Scalable Collectives
06ARCHITECTURE

Scientific Software Engineering

Modular Software Design

Applying modern C++ paradigms, zero-cost abstractions, automated regression, and Continuous V&V to sustain long-lived simulation codes.

  • Modern C++17
  • CMake / CTest
  • RAII
  • V&V Suites
07MICRO-ARCHITECTURE

Performance Optimization

Cache, SIMD & Roofline

Analyzing cache line utilization, memory bandwidth saturation, and SIMD vectorization to approach theoretical roofline hardware performance limits.

  • Roofline Model
  • SIMD / AVX-512
  • Cache Alignment
08PORTABILITY

Performance Portability

Single-Source Dispatch

Writing portable mathematical kernels that target multi-vendor CPUs and accelerators through modern unified abstraction frameworks.

  • Kokkos
  • C++ Parallel Algorithms
  • Hardware Targets

TECHNICAL METHODOLOGIES & RUNTIMES

  • Modern C++
  • MPI / OpenMP
  • Kokkos / CUDA
  • SLURM / CMake
  • Numerical methods
  • Software architecture

BENCHMARKING / EVIDENCE

Measure the system before optimizing it.

Benchmarking is treated as a reproducible loop: define the workload, capture a baseline, compare a change, and keep the evidence attached to the result.

REPEATABLETRACEABLECOMPARABLE
PERFORMANCE LABNO SYNTHETIC RESULTS
01
WORKLOADdefine input and environment
SET
02
BASELINEcapture the reference run
CAPTURE
03
COMPAREinspect traces and behavior
ANALYZE
04
EVIDENCErecord what changed
REPORT