Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
結論:IBM Granite 4.0は、最大モデルの性能競争よりも、企業のRAG、エージェント、長文処理、オンプレミス推論で、性能・推論コスト・監査可能性のバランスを重視する組織に適したオープンウェイトLLMです。
ただし、IBMが示す「メモリ使用量70%以上削減」「推論速度2倍」は特定条件での比較結果であり、すべての環境に当てはまる保証ではありません。また、2026年4月にはGranite 4.1も発表されています。新規導入では、4.0を選ぶ理由をランタイム対応、既存検証、契約、運用コストなど具体的に確認すべきです。
IBM Granite 4.0とは
Granite 4.0は、IBMが2025年10月2日に発表したオープンウェイト言語モデル群です。従来型Transformerだけでなく、Mamba-2とTransformerを組み合わせたハイブリッド構成を採用し、一部のモデルではMixture of Experts(MoE)も使います。
Recommended Free Tools
より正確には「オープンソースLLM」ではなくオープンウェイトモデルです。モデルの重みが公開され、Apache 2.0ライセンスで利用しやすい一方、学習コード、学習データ、推論ランタイム、再配布条件がすべて同じ範囲で公開されるとは限りません。ウェイトを無料で取得できることと、推論や運用が無料であることも別です。
#1 Best Overall
- 100% Satisfaction Warranty – Our servers book for waitress organization are handcrafted with elegant stitching that lasts. We take pride in offering our customers a waitress book made to exceptional quality standards. To ensure satisfaction, every waiters checkbook is backed by a 1-YEAR WARRANTY. If you are not 100% SATISFIED for any reason we will send you a replacement. No Questions Asked
- Holds up under Pressure – When you're taking orders the last thing you need is a flimsy waiter book that keeps bending. Our 8”x5” server books for waitress organization is the only one with a premium reinforced dual inner core. Providing an unmatched sturdy reliable writing surface that will last for years
- On Another Level – Halt the endless cycle of replacing your cheap thin black server book that barely lasts a week. This serving book for waitresses can become your permanent partner. Crafted with overwhelmingly strong attention to detail, the waiter checkbook offers an unparalleled value that you won’t regret investing in
- Scribble In Style – Impression is everything. You’re making a statement when you bring out this sleek vegan leather serving book. Our serving books have no logos or images and exquisite stitching for a professional feel your colleagues will envy
- Stay Calm and Collected – Whether you have 1 table or 7, organization is key. This server checkbook has 9 versatile pockets including a durable metal zipper to keep your cash secure. Stay on top of everything with this deluxe server book organizer and bring superior service to every customer
Granite 4.0は汎用チャットだけでなく、企業向けRAG、関数呼び出し、エージェント、構造化出力、長文文書処理、エッジ推論を想定しています。Baseモデルは追加学習や独自の調整を前提とし、Instructモデルは指示追従や対話、ツール利用に向きます。モデル一覧と用途はIBMの公式モデル資料で確認できます。
モデル構成とサイズ
| モデル | 構成 | パラメータ | 向く用途 |
|---|---|---|---|
| Granite-4.0-H-Small | ハイブリッドMoE | 32B総量/9B活性 | 企業RAG、エージェント |
| Granite-4.0-H-Tiny | ハイブリッドMoE | 7B総量/1B活性 | 低遅延、ローカル推論 |
| Granite-4.0-H-Micro | ハイブリッドDense | 3B | 分類、抽出、関数呼び出し |
| Granite-4.0-Micro | 従来型Dense | 3B | 既存ランタイムとの互換性 |
| Granite-4.0-H-1B | ハイブリッドDense | 1.5B | エッジ、端末、低遅延 |
| Granite-4.0-1B | 従来型Dense | 1B | llama.cpp、PEFTなど |
| Granite-4.0-H-350M | ハイブリッドDense | 350M | 最小構成、軽量推論 |
| Granite-4.0-350M | 従来型Dense | 350M | 互換性重視の軽量用途 |
MoEでは、総パラメータ数と1トークン生成時に計算する活性パラメータ数が異なります。H-Smallは32Bの重みを持ちながら、活性パラメータは9Bです。そのため「9Bモデル」と単純化するのも、「32B相当の計算量」と説明するのも不正確です。ロード時のメモリ、通信、キャッシュ、ランタイム実装では総量と活性量の両方を確認する必要があります。
Mamba-2とTransformerのハイブリッド構成
Transformerはトークン間の関係を柔軟に扱いやすい一方、長いコンテキストや複数の同時セッションではKVキャッシュのメモリ負荷が大きくなりやすい構造です。Mamba系の状態空間モデルは、長い系列をより効率的に処理できる可能性があります。
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Granite 4.0は両者を組み合わせ、長文処理と指示追従、ツール利用の両立を狙います。IBMは、同等クラスのモデルと比べ、特に長文・複数セッション推論でメモリ使用量を70%以上削減し、推論速度を2倍にできると説明しています。ただし、これはIBMの比較条件に基づく主張です。GPU、量子化、ランタイム、バッチサイズ、コンテキスト長で結果は変わります。詳細はIBMの発表資料を参照してください。
ハイブリッドモデルは、すべての推論環境で通常のTransformerモデルと同じように動くわけではありません。Transformers、vLLM、SGLang、llama.cpp、NVIDIA NIMなど、利用予定のランタイムが対象モデルのMamba、MoE、量子化に対応しているかを事前確認する必要があります。
Rank #2
- 【Perfectly Fit in Server Aprons】: Our black server book size is 8.15" x 5.12" x 0.59", which can hold a regular guest checkbook and is handy to be carried in a server apron pocket, won’t be too tight or too big, efficiency as a server money holder.
- 【Stay Organized All in Needs】: 9 compartments and 1 pen holder in one serving book, with a zipper pocket to store your coins, changes, and money. Multi-functional pockets to organize checkbooks, cash, ticket books, server pads, credit cards, coupons, or any other paper documents, nice waitress accessories partner for servers.
- 【Waterproof Leather Material】: The waitress book is made of premium sturdy PU leather, Eco-friendly and odorless, features excellent workmanship and tight stitching, easy to clean. Plus an elastic pen loop to be a nice waitstaff organizer to help you hold the pen that is always away from home and improve the service speed.
- 【Portable and Long-lasting】: Our server books for the waiter are lightweight to carry around, and sturdy as a guest checkbook holder, premium material makes them sturdy and won’t easily deform or press the belly when bent over.
- 【100% Satisfaction Guarantee】: We hope you love your server book wallet and place your order with confidence, all of our men’s & women’s server books are backed by a replacement guarantee. Any questions will be answered within 24 hours.
最大コンテキスト長は131,072トークンとされています。しかし、128Kトークンを入力できることは、128Kを現実的な速度と費用で処理できることや、文書中の情報を正確に使えることを意味しません。情報位置による見落とし、検索ノイズ、指示の競合、出力コストを自社文書で評価すべきです。
企業ワークロードで見る性能
Granite 4.0を評価するとき、一般的な知識ベンチマークだけでは不十分です。企業導入では次の指標を測定します。
- RAGで検索結果に基づいて回答できるか
- 引用や根拠を正確に提示できるか
- JSONなどの構造化出力を守れるか
- 適切なツールを選び、正しい引数を生成できるか
- 複数ターンの状態を維持できるか
- 長文文書から必要な項目を抽出できるか
- 同時実行時のp50、p95、p99レイテンシーとスループット
- 失敗時の再試行、拒否、フォールバックが適切か
IBMはGranite 4.0 Smallについて、指示追従や関数呼び出しなどエージェント系タスクで強い結果を示しています。ただし、確認できる主要資料はIBMによる公式発表とモデル資料が中心です。「業界最高」や「GPT級」といった断定ではなく、IBMが報告するベンチマーク結果として扱うのが適切です。
RAGと関数呼び出しの実務テスト
長文RAGでは、単にコンテキストを増やすのではなく、チャンク分割、再ランキング、引用検証、回答拒否条件を設計します。日本語の契約書、稟議書、製品マニュアル、表データ、OCR誤り、社内略語を含む評価セットも必要です。英語中心のベンチマーク結果を、日本語業務の精度にそのまま外挿してはいけません。
関数呼び出しでは、必須引数の欠落、不正な引数、列挙値の誤り、ツール停止、タイムアウト、同じ操作の二重実行、破壊的操作の承認漏れをテストします。モデルがベンチマークで良い結果を出していても、実際のAPIスキーマと認証・承認フローでは失敗する可能性があります。
Rank #3
- Adequate quantity: we have prepared 6 pieces of server books with zipper pocket in the package, sufficient quantity can easily satisfy your daily use and replacement requirements, making your work more efficient and convenient
- Abundant capacity: with 8 pockets design, including the credit card holder, window viewer, receipt pocket, vertical zipper pocket, order pad holder sleeve and pen holder, this waiter book can help you organize items separately and methodically
- Fine workmanship: our serving book is made of quality PU leather, with a protective clear coating layer, sturdy and reliable, not easy to stain, tear or fade, smooth on surface, providing you with a nice use experience, and can serve you for a long time
- Proper size and portable: each black server book measures around 8.07 x 4.92 x 0.39 inches in closure size, and its expansion size is around 10.35 x 4.92 inches, a suitable size for most people, and you can put it in your pocket for use
- Versatile applications: this server wallet can be widely adopted for serving, cleaning, gardening, cooking, baking, crafting and more; In addition, it can hold various small tools, such as check pads, napkins, cards, pens, recipe cards, menus and so on
「安い」の意味を4つに分ける
Granite 4.0のコスト優位性は、単価だけで判断できません。少なくとも次の4項目を分けて計算します。
Free tools Windows power users keep installed
One-click scans. No signup required.
- モデル利用料:APIやマネージド推論の入力・出力トークン料金。
- 計算基盤費:GPU、CPU、ストレージ、電力、冷却、ネットワーク。
- エンジニアリング費:ランタイム対応、量子化、監視、更新、障害対応。
- ガバナンス費:監査、ログ保存、アクセス管理、データ分離、評価、契約対応。
watsonx.aiの公開価格を使った計算例
2026年8月時点でIBMの米ドル公開価格ページに表示されている参考値では、Granite-4-H-Smallは入力100万トークンあたり0.06ドル、出力100万トークンあたり0.25ドルです。地域、税、契約、プラン、提供状況で変わるため、導入前に最新の料金ページを確認してください。
たとえば月間入力が10億トークン、出力が2億トークンなら、単純計算は次の通りです。
- 入力:1,000百万×0.06ドル=60ドル
- 出力:200百万×0.25ドル=50ドル
- モデル推論料金の合計:約110ドル
これはモデル推論料金だけの概算です。プラン料金、RAGの検索基盤、埋め込み、ストレージ、監視、データ転送は含みません。IBMのドキュメントには、入力1,000トークンあたり0.0000636ドル、出力1,000トークンあたり0.000265ドルという表示もあります。表示単位や価格は更新されるため、比較時は100万トークン単位に統一してください。
メモリが減っても総保有コストが必ず下がるわけではありません。MambaやMoE対応カーネルの保守、量子化後の品質確認、低いGPU稼働率、監視や更新の人件費が大きければ、単純なGPU削減効果は相殺されます。
Rank #4
企業ガバナンスとセキュリティ
評価できる要素
- Apache 2.0による商用利用、改変、再配布の柔軟性
- 暗号署名されたモデルチェックポイントによる出所・改ざん確認
- IBMが説明するISO/IEC 42001認証
- モデルカードや学習データに関する情報
- AIの構築、学習、検証、デプロイを記録するAI bill of materialsの考え方
- watsonx.ai経由のIBMモデルに契約上の補償が適用される場合があること
関連するIBMの説明はGranite 4.0の発表、AI bill of materialsの解説、基盤モデルの契約情報で確認できます。
過大評価してはいけない点
ISO/IEC 42001はAIマネジメントシステムに関する認証であり、モデルの回答が常に正確、安全、無害であることの証明ではありません。自社データ、プロンプト、ユーザー権限、監視、ログ、リスク管理まで自動的に保証するものでもありません。
Apache 2.0でも、学習データに由来する著作権や個人情報、輸出規制、業界規制の確認は必要です。自社でウェイトを取得して量子化・再配布すれば、元の署名やサプライチェーン管理との関係も確認しなければなりません。署名済みモデルであることと、自社業務に安全であることも別です。
モデルサイズの選び方
| 候補 | 選びやすいケース | 注意点 |
|---|---|---|
| H-Small | 本番RAG、エージェント、128Kコンテキスト、GPU運用 | MoEとMamba対応ランタイムが必要 |
| H-Tiny | 低遅延、長い入力、ローカルGPU、補助モデル | 品質と同時実行数を自社測定 |
| H-Micro/Micro | 分類、ルーティング、抽出、関数呼び出し | 回答生成より前処理向き。Microは互換性重視 |
| 1B/350M | 端末、組み込み、単純な分類、オフライン処理 | 複雑な推論や長文回答には不向きになりやすい |
導入経路の比較
watsonx.ai
最短で企業検証を始めるならwatsonx.aiが候補です。無料枠、Essentials、Standardのプランがあり、公開ページではStandardが月額1,110ドルからと表示されています。これはプラットフォーム料金であり、モデルのトークン料金とは別に考えます。RAG、エージェント、評価、ガバナンス、企業サポートをまとめやすい一方、完全オンプレミスや既存GPUの活用を優先するケースには向きません。
watsonx.aiの製品ページと料金ページで契約条件を確認してください。
Best Value
- Stay Organized: Our 1-part bond paper checkbooks are perfect for any server! Sized to fit perfectly in standard apron pockets, our waitress notepad will help you record vital information for serving and accounting purposes.
- Record Information: With room for writing the date, table and number of guests, you'll never lose track of your orders. Our waiter book also has prompts for appetizers, soup or salad, entrees, vegetables or potatoes, dessert, and beverages.
- Detachable Receipt: Designed with a perforated guest receipt that you can give as a customer copy or keep for record keeping. We've provided extra rows on the back for additional note taking.
- Numbered Checks: Each ticket has a unique serial number printed at the top, helping prevent numerical and clerical errors. Great for restaurant, diner, café, and food truck orders.
- Great Value Pack: Includes 5 pink server note pads. Each booklet has 50 bound order slips - that's 250 ticket sheets total! Pad measures: 6.75” L x 3.5” in W. Perforated stub measures: 3.5” L x 0.75” W.ted stub measures 3.5” Length x 0.75” Width
Hugging Face
Hugging Faceではモデルウェイトを取得でき、Inference ProvidersやInference Endpointsでマネージド推論も利用できます。モデルのリビジョンや推論環境を柔軟に管理しやすい一方、IBMの契約上の補償や統合された企業ガバナンスを同じ形で得られるとは限りません。料金はHugging Faceの課金説明を確認してください。
ローカル・自社GPU
機密データを外部に出せない、オフライン環境で使う、長期的に大量推論を行う場合に適します。費用の中心はウェイトではなく、GPU、電力、保守、MLOps、監視、人員です。Hugging Faceから取得してTransformersで検証する例は次のようになりますが、モデルIDやライブラリ対応は更新され得るため、実行前に公式リポジトリを確認してください。
pip install torch transformers accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "ibm-granite/granite-4.0-h-tiny"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto"
)
prompt = "企業向けRAGシステムで回答の信頼性を高める方法を3つ説明してください。"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Ollama、LM Studio、NVIDIA NIMなど
IBMはGranite 4.0の提供先として、Ollama、LM Studio、NVIDIA NIM、Docker Hub、Kaggle、Replicateなどを挙げています。ただし、サービスごとに使えるモデル、量子化、料金、リージョン、商用条件は異なります。「どこでも同じGranite 4.0が同じ性能で動く」とは考えず、パートナー実行環境の資料で対応状況を確認します。NVIDIA GPU基盤の標準化ならNVIDIA NIM、OpenShift中心の組織ならRed Hat系の提供状況も候補になります。
Granite 4.0をGranite 4.1や他モデルと比較する
2026年4月にはGranite 4.1が発表されました。IBM Researchは、Granite 4.1 8B InstructがGranite 4.0の32B MoEモデルと同等以上の性能を示すケースがあると報告しています。したがって、4.0をIBMの最新モデルと表現することはできません。新規案件ではGranite 4.1を必ず比較対象に含めます。
それでも4.0を選ぶ合理性は、既存の評価済みモデルであること、特定ランタイムへの対応、MoEによる処理効率、既存契約や運用基盤との整合性にあります。反対に、これから新規開発を始めるなら、4.1の性能、ファインチューニング、運用柔軟性を検証しない理由はありません。
Meta Llamaはエコシステム、ツール、ホスティング選択肢の広さが強みになりやすく、Mistralは小型・中型モデルや効率的な推論を比較する候補です。ライセンス、コンテキスト長、ツール呼び出し、量子化、商用条件はモデルごとに異なります。watsonx.aiではGranite以外にLlamaやMistralも扱えるため、同一基盤で自社評価しやすい場合があります。
本番導入前のチェックリスト
- RAG、分類、抽出、コード、関数呼び出しなど対象タスクを明確にする。
- 日本語の実データに近い評価セットを作り、正解率、引用精度、拒否精度を測る。
- p50、p95、p99レイテンシー、初回トークン、出力速度、同時実行数を測る。
- 重み、KVキャッシュ、MoE総量、活性パラメータ、量子化後のメモリを実測する。
- Transformers、vLLM、SGLang、llama.cpp、NVIDIA NIMなどの対応版を固定する。
- JSON遵守率、引数エラー、タイムアウト、二重実行、破壊的操作の承認をテストする。
- モデルリポジトリ、リビジョン、チェックサム、量子化方式を記録する。
- システムプロンプト、安全フィルター、評価データ、合格基準を保存する。
- 入力データ、ログ、保持期間、アクセス権限、学習利用の有無を契約と運用で確認する。
- 更新時の回帰テスト、ロールバック、監査ログ、障害時のフォールバックを準備する。
最終判断
Granite 4.0は、企業の小型・中型ワークロード、RAG、エージェント、長文処理、ローカル推論で有力な選択肢です。Apache 2.0、署名付きチェックポイント、IBMが説明するISO/IEC 42001や企業向け提供経路は、導入時のライセンス・出所・ガバナンスを整理する材料になります。
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors一方で、これらは回答品質や安全性を自動的に保証しません。IBMの70%削減、2倍高速という主張も、実際のGPU、ランタイム、量子化、同時実行、コンテキスト長で再現性を確認する必要があります。2026年時点の新規案件ではGranite 4.1、Llama、Mistral、自社ホスティングを含め、モデル料金だけでなくGPU稼働率、再試行率、有人確認率、ツール失敗率、監査費まで含めて比較することが採用判断の核心です。
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

