65日以内にGPUが必要ですか?私たちにお任せください

私たちにおまかせください

ミドクラジャパン株式会社、誰でも簡単にAI技術を作成したり、利用したりできるよう、ネットワークやインフラストラクチャーをシームレスに統合し、スケーラブルな最先端ソリューションをご提供することで、世界中の企業を支援しています。

我々のミッションは、イノベーションと現実世界のアプリケーション間ギャップを埋め、あらゆる規模の企業の方々がAI技術を効果的に活用できるようにすることです。高度な専門知識を駆使し、常に進化し続ける技術にコミットすることで、お客様がより効率的に、的確な意思決定をできるよう支援します。AIを活用した産業の未来を創造し、形にしていきます。

ミッションを達成するための3本柱

お客様のご都合に合わせて、迅速に納品いたします。
もし具体的に時期がお決まりでしたら、お気軽にお問い合わせください。通常プロセスより短期で納品できるかもしれません。

パフォーマンス

  • お客様の環境に適したGPUサーバーを設定します
  • すべてのリソースを効率的に使えるよう最適化します
  • 高スループットと低レイテンシーを実現します

信頼性

  • ハードウェア障害頻度を低減します
  • お客様のご要望にできる限りお答えします
  • お客様にご満足いただけるサービスレベルを遵守します

スケーラビリティ

  • GPUクラスタ環境を簡単に利用できるようにし、お客様は学習/推論モデルにフォーカスできるようにします

テスト項目

クラウドの構築し運用開始するために必須な8つの検証を実施※テスト結果は、ご購入者様の社内利用で利用可能。
社外への共有については、別途ご相談ください

Hardware Testing

Test Items

Components Verification

  • GPUs, CPUs, Memory, and Network Cards

​Thermal and Power Test

  • Ensure the systems meets operation temperature range ((5°C – 30°C)

Test Tool

  1. Nvidia NVSM(Nvidia System Management): Monitors Nvidia DGX/HGX system by providing insight into system temperatures, power usage and overall components health
  2. Stress-ng: Stress CPUs, Memory, and other hardware components under heavy load to detect potential hardware failures or thermal issues
  3. HPL(Linkpack): A benchmark tool to stress the GPUs and CPU

System Testing

Test Items

Optimize Configuration

  • BIOS settings, Driver install

​Benchmark Performance

  • MLPerft and SPEC

Test Tool

  1. MLPerf: benchmarking tool to measure the performance of machine learning and deep learning models
  2. Linkpack: benchmark CPUs, Memory, and compiler performance

Network Testing

Test Items

Ethernet and IB​

  • Run additional RDMA benchmarks to confirm low-latency communication across the IB

Fabric

  • Front-end and RDMA/back-end Network – front-end (system operation) and RDMA/back-end (System health monitor) are properly segregated to ensure security and performance.

Test Tool

  1. IPERF: measure network bandwidth and quality
  2. NetPerf: benchmark the throughput and latency in networked system over Ethernet and RDMA setups
  3. Nvidia HPC-X: test network performance in multi-nodes setups

CUDA Testing

Test Items

CUDA Compatibility​

​Multi-GPU tests

  • Ensure the multi-GPU communication is function properly

Test Tool

  1. CUDA Toolkit Samples: Verify the correct installation and configuration of CUDA
  2. Nvdia-smi: monitor the GPU health, usage, and performance
  3. Nvbandwidth: specific test NVLink bandwidth between GPUs

Storage Testing

Test Items

Storage Fabric performance

  • Simulate heavy read and write operations to make sure throughput exceeding 40GBps per nodes

​Latency and IOPS

Test Tool

  1. FIO (Flexible I/O Tester): simulate different I/O workloads to test the storage performance
  2. Nvidia HPC-X: benchmark parallel I/O operations over a networked storage solution

AI Training Testing

Test Items

LLaMA Model Training

  • Focuses on training throughput, model convergence speed, and ability to scale to more  GPU

Performance Monitor

  • Monitor GPU utilization, memory usage, and overall training time

Test Tool

  1. LLaMA Model (Meta AI): large datasets to benchmark the system AI training capabilities
  2. TensorFlow/PyTorch: ideal for Multi-GPU tests
  3. TensorBoard: monitor training performance metrics such as GPU utilization, memory, consumption, and the rate of convergence during LLaMA training

HPC Workload Testing

Test Items

cuFFT (CUDA Fast Fourier Transform)

  • Evaluate the performance of FFT operations on GPUs

cuBLAS (CUDA Basic Linear Algebra Subroutines)

  • Validates the efficiency of numerical computation across multiple GPUs

Test Tool

  1. cuFFT: perform fast Fourier transforms (FFTs) on large datasets
  2. cuBLAS: focus on matrix multiplications and other linear algebra operations using cuBLAS
  3. Nvprof/Nsight system: measuring the performance of cuFFT and cuBLAS