SaaS API rate limiting: tenant quotas, 429 responses और fair use
Tenant-aware API rate limits और quotas बनाएँ, उपयोगी 429 responses दें, shared resources बचाएँ और असली request cost के अनुसार limits test करें।
इस मार्गदर्शिका में
SaaS में API rate limit और quota क्या हैं?
Rate limit कम अवधि या burst में caller कितनी requests भेज सकता है, इसे नियंत्रित करती है। Quota लंबी अवधि का उपयोग सीमित करती है, जैसे हर दिन exports की संख्या या monthly API allocation। SaaS teams इन्हें capacity बचाने, abusive traffic सीमित करने और shared service को एक tenant से प्रभावित होने से रोकने के लिए रखती हैं। Limits केवल request count नहीं, operation की लागत और उद्देश्य देखकर तय करें।
Caller की पहचान किससे होगी, यह चुनें
जोखिम के अनुसार authenticated tenant, API credential, user या इनके मेल पर limit लगाएँ। केवल IP की limit साझा network पर कई वैध users को प्रभावित कर सकती है, जबकि केवल credential की limit से कई keys बनाकर compromised account गतिविधि कर सकता है। Authentication तथा tenant authorisation को counting mechanism से अलग रखें।
Burst, request cost और concurrency का हिसाब रखें
Token bucket average rate सीमित करते हुए थोड़े समय का नियंत्रित burst दे सकता है। महँगी query, बड़े payload या लंबे job के लिए weighted cost, maximum concurrency, queue limit या total quota भी सोचें। समान requests-per-second वाले दो endpoints database या compute की अलग मात्रा ले सकते हैं।
Limit पर caller को स्पष्ट response दें
HTTP 429 बताता है कि caller ने एक अवधि में बहुत requests भेजीं। उपयोगी ढंग से limit समझाएँ और सम्भव हो तो Retry-After value दें ताकि client अगली कोशिश से पहले इंतज़ार करे। Internal capacity या दूसरे tenant का usage उजागर न करें। HTTP standard यह तय नहीं करता कि server caller की पहचान या request count कैसे करेगा।
| Operation / resource cost | Caller identity | Rate / burst | Quota या concurrency cap | 429 और retry व्यवहार |
|---|---|---|---|---|
| Read API | ||||
| महँगी report या export | ||||
| Write या asynchronous job |
SaaS team API limits कैसे तय करे?
पहले सामान्य और peak demand मापें
Traffic तथा load-test evidence से supported request volume, burst size, payload size और downstream cost का अनुमान लें। वैध customer workflow और dependencies के लिए headroom रखते हुए शुरुआती सीमा तय करें। Tested limit दर्ज करें और data volume, infrastructure या customer plans बदलने पर review करें।
एक से अधिक shared layer पर नियंत्रण रखें
Edge limit application तक flood पहुँचने से रोक सकती है, लेकिन मंजूर हुई request queue, worker, database या third-party capacity ले सकती है। System के भीतर उपयुक्त concurrency तथा resource limits भी लगाएँ। Async काम में tenant identity साथ रखें, ताकि queued jobs तय budget से बचकर न चलें।
Limits को customer plan के साथ स्पष्ट और एक-जैसा रखें
हर plan के लिए unit, समय अवधि, burst व्यवहार, reset तरीका, fair-use terms और support route लिखें। Product copy, API docs, dashboards और enforcement में वही values रखें। “Unlimited” कहकर छिपी restrictive cap न लगाएँ; customer-facing terms का contract review करें।
API client throttled response पर क्या करे?
दोबारा कोशिश से पहले रुकें और Retry-After मानें
Rejected request को client तुरंत फिर न भेजे। Response में Retry-After हो तो कम से कम उतनी देर रुके; header न हो तो सीमित backoff लगाएँ और बहुत से clients एक साथ retry न करें इसके लिए jitter जोड़ें। Retries की सीमा रखें और operation पूरा न हो तो उपयोगी error दिखाएँ।
Retry होने वाले writes को सुरक्षित बनाएँ
Network timeout के बाद client को पता न हो कि write सफल हुआ या नहीं। Retry होने वाली operations में idempotency key या दूसरा duplicate protection रखें और throttled response को अनिश्चित result से अलग समझें। Payment या data mutation को अंधाधुंध दोबारा न भेजें।
स्वीकृत traffic के साथ rejected requests भी monitor करें
Privacy बनाए रखते हुए endpoint, tenant plan और caller type के अनुसार 429 rate देखें। अचानक वृद्धि client retry bug, गलत limit या abuse का संकेत हो सकती है। Legitimate workflow वाले rejected users भी देखें और safeguards हटाए बिना evidence से policy सुधारें।
SaaS API rate limits पर सवाल
Rate limit और quota में क्या फर्क है?
Rate limit कम समय में request की गति या burst नियंत्रित करती है। Quota लंबी अवधि में कुल उपयोग सीमित करती है। मजबूत API को दोनों के साथ महँगी operations के लिए cost तथा concurrency limits चाहिए हो सकती हैं।
क्या केवल IP address के आधार पर rate limit लगानी चाहिए?
आमतौर पर इसे अकेला control न बनाएँ। एक IP के पीछे कई users हो सकते हैं और एक user का IP बदल सकता है। API तथा abuse risk के अनुसार network protection को authenticated tenant, credential या user limits के साथ रखें।
क्या हर 429 response में Retry-After होना चाहिए?
RFC 6585 कहता है कि 429 response में Retry-After दिया जा सकता है। जब server सुरक्षित wait time बता सकता हो तो यह उपयोगी है। Header न होने पर client क्या करे, यह दस्तावेज़ करें; system जितना सह न सके उससे छोटा retry समय न सुझाएँ।
क्या API rate limit से यह guarantee होती है कि एक tenant दूसरे को प्रभावित नहीं करेगा?
नहीं। Limits resource exhaustion घटा सकती हैं, लेकिन isolation downstream concurrency, queue capacity, database व्यवहार, monitoring और tested architecture पर भी निर्भर है। Cross-tenant impact जाँचने के लिए tenant-aware load tests चलाएँ।
संबंधित व्यावहारिक मार्गदर्शिकाएँ
संबंधित मुद्दों की मार्गदर्शिकाएँ
स्रोत और प्रकाशन रिकॉर्ड
Draft prepared 27 September 2026; project-team editorial review pending · स्रोत जाँचे गए .