One self-hosted console to run your entire business — commerce, ERP, HRM, CRM & manufacturing

भाग 4: Valkey, Queue और Log, नीरब अंश जो पहले टूटते हैं

NovaCommerce case study part 4: 2,854 Valkey timeout, एक worker जो अभी भी single point of failure है, 115 मिनट delayed merit card, AI provider ने "no" कहा, और building को गिनने वाला rate limiter।

Author

Anichur Rahaman

एक दिन पहले13 min read3 views
भाग 4: Valkey, Queue और Log, नीरब अंश जो पहले टूटते हैं

6 अक्टूबर को 02:28 पर एक backend load test 2,854 application error दे गया। हर एक same sentence था, throw करते समय connecting: RedisException: Operation timed out। Web server healthy थे और database healthy था। Part जो "no" कहता था वह quiet service था session, cache और queue hold करता।

Same week ने एक दूसरी संख्या produce किया: 115। यह कितने मिनट slowest merit card एक exam के बाद recompute होने का wait किया। Worker का CPU cause नहीं था, और WebSocket भी नहीं। Job सिर्फ़ गलत line में खड़ा था।

Part 3 database को cover किया। यह part उन piece को cover करता है जो एक stateless fleet को tie करते हैं: Valkey, एक fixed worker, queue, rate limit, log और signal जिन्हें हम watch करते हैं। वे एक diagram पर dull दिखते हैं, और उन्होंने हमें हमारे real surprise के बहुत दिए। नीचे का हर section एक failure, एक cause और एक change है।

यह "From One Server to Exam-Day Ready" five-part case study का हिस्सा 4 है। NovaCommerce एक काल्पनिक नाम है; architecture, संख्याएँ और गलतियाँ वास्तविक हैं।

एक fleet जिसके पास कोई memory नहीं है कहीं याद रखने के लिए

एक node जो किसी भी minute पर delete किया जा सकता है कुछ own नहीं कर सकता। इस platform में owner एक छोटी list हैं: managed MySQL application data के लिए, managed PostgreSQL pgvector के साथ AI embedding के लिए, Spaces object storage file के लिए, managed OpenSearch log के लिए, और managed Valkey हर चीज़ के लिए छोटी, fast और shared: session, cache और queue।

Supermarket cashier की तरह सोचो। Till drawer और stock list एक back office में रहते हैं। एक cashier mid-shift पर घर जा सकता है और अगला lane ले सकता है, क्योंकि cashier की pocket में कुछ नहीं था।

Layered diagram: frontend pool, backend pool और fixed worker एक private network line के ऊपर, और नीचे पाँच store: Valkey, MySQL, PostgreSQL pgvector के साथ, Spaces और OpenSearch
ऊपर node कुछ नहीं hold करते। हर चीज़ जो एक node deleted होने से survive करना चाहिए नीचे store में से एक में रहती है।

हमने assume करने की जगह यह check किया। 7 अक्टूबर पर backend node ने 24 घंटे में 0 file local disk में लिखे। Upload, written-answer PDF सहित, सीधे Spaces को जाते हैं; session और cache Valkey में हैं; log stdout जाते हैं। Implemented और verified। यही वजह है कि autoscale pool किसी भी node को कभी भी delete कर सकते हैं।

Stateयह कहाँ रहता हैयह हमें क्या खरीदता है
SessionValkey (primary + standby)कोई भी backend node कोई विद्यार्थी serve कर सकता है
CacheValkeyएक cache हर node से share किया
Queued jobValkeySubmission safely wait करते हैं अगर worker down हो
Upload, answer PDF, mediaSpaces CDN के साथकोई file कभी node disk पर बैठता नहीं
Logstdout, OpenSearch को shippedवे node को outlive करते हैं (backend done, frontend planned)
Application dataManaged MySQLPart 3 में covered

एक Valkey, purpose पर

Cache, session और queue एक managed Valkey 8 एक standby node के साथ share करते हैं। इसकी eviction policy noeviction है: जब memory full है, Valkey नए write से noeviction के साथ error देता है quietly old key delete करने की जगह। एक pure cache के लिए यह backwards दिखता है। लेकिन यह instance session और queued job भी hold करता है, और हमें एक loud error prefer हैं एक job vanish होने के बिना trace। हमें उस edge से far हैं: exam evening 7 अक्टूबर पर यह लगभग 12.5% अपनी memory use करता है।

2,854 timeout: जब shared memory answer नहीं दे सके

02:28 पर वापस। Test एक 4 GB Valkey को hit कर रहा था primary और standby के साथ। Memory problem नहीं था। Load के अंतर्गत यह नए connection accept fast enough नहीं कर सकता, और हर request जो 5-second connect timeout से ज़्यादा wait करता एक 500 बन जाता।

दबाव का कुछ हमारा था। हर request Valkey को fresh TLS connection खोलता, जिसे हमने लगभग 5.4 ms CPU per request पर measure किया (MySQL के लिए लगभग 3.3 ms के विपरीत)। 65 to 70 requests per second एक backend node serve कर सकता है पर, वह handshake अकेले लगभग एक third का एक core cost करता है। यह back-of-the-envelope arithmetic है, profile नहीं।

Fix एक resize था, in place किया: Valkey 8 GB, दो node (primary plus standby) में गया, data kept, hostname unchanged। Application change नहीं, cutover नहीं। Implemented।

फिर हमने फिर से measure किया। Tested: 25 से 500 requests per second तक एक gradual climb, 2 backend node पर शुरू करते हुए pool scale out, 0 application error और 0 backend 5xx दिए। Load balancer 143 error show किया, लेकिन सिर्फ़ जबकि node CPU-saturated थे: expected "out of capacity" signal, bug नहीं।

Useful lesson यह है कि bottleneck कुछ दिनों में तीन बार moved, और हर move को एक अलग kind का fix चाहिए।

OrderBottleneckProblem का kindFix
1API rate limiter IP से keyedCode और configurationAccount से key करो (Implemented)
2Valkey connectionData-tier sizing4 GB to 8 GB, 2 node (Implemented)
3Backend CPU per nodePlain capacityज़्यादा node, pre-scaled (Implemented)

सिर्फ़ अंतिम row ज़्यादा server से solve होता है। दूसरे के लिए, ज़्यादा server अपने connection को same Valkey में open करते।

एक worker, और काम सिर्फ़ यह कर सकता है

Worker के चारों ओर हर चीज़ scale करती है। Worker अपने आप नहीं। यह एक fixed droplet है (4 vCPU, 8 GB) और यह चार काम carry करता है:

  • Queue। Default, exam, तीन notification queue, OMR और enrollment report, Horizon द्वारा run किए जाते हैं।
  • Scheduler। Timed task बिल्कुल एक बार run करने चाहिए।
  • WebSocket। Reverb live update serve करता है। इसका port सिर्फ़ backend node को open है, tag से।
  • सब SMS। SMS gateway एक single IP address को whitelist करता है, तो SMS सिर्फ़ इस machine से leave कर सकता है।

यही वजह है कि queue, scheduler, WebSocket और OMR web node पर switch off हैं: एक नया autoscale node कभी एक काम दो बार नहीं चलाना चाहिए। जो हमारा case add करता है generic rule के लिए यह है कि SMS एक process से नहीं, एक address से tie है।

यह भी एक single point of failure है, और हम plainly कहते हैं। अगर यह droplet मर जाता, SMS, scheduler और exam processing रुक जाता। विद्यार्थी submission जारी रखते हैं, और उनके submission safely Valkey में wait करते हैं, जो reason है queue वहाँ रहते हैं। लेकिन कुछ भी उन्हें process नहीं करता जब तक worker back न हो।

7 अक्टूबर की exam evening पर worker 46% CPU पर peaked 0 failed job के साथ, तो यह एक risk है हम plan करते हैं, failure नहीं हमने देखा। Plan, honest label के साथ:

Stepक्योंStatus
Reserved IP, SMS gateway के साथ whitelist कियाReplacement worker एक known address से SMS भेजता हैPlanned
छोटा standby workerकुछ queue और scheduler पर take over करने के लिए तैयारPlanned
Exam-only burst worker image, बड़े exam से पहले start कियाExtra queue capacity जब इसकी जरूरत हैPlanned
Worker का एक snapshotRebuild एक image से शुरू होता है, memory नहींImplemented

Reserved IP पहले आता है: इसके बिना, एक standby worker एक address से SMS भेजता जिसे gateway accept नहीं करता।

Merit card जो 115 मिनट wait किया

एक exam के बाद, "merit recompute" job हर विद्यार्थी का merit card refresh करता है। 6 अक्टूबर पर वे card 10 to 115 मिनट late appear होने लगे।

हमने क्या rule out किया

Obvious suspect worker का CPU और WebSocket server थे। दोनों cause नहीं थे।

यह क्या था

Merit job tiny है: 0.34 second। यह एक queue share करता था, 2 worker से serve किया गया, एक AI job के साथ जो लगभग 10 second एक language-model call बनाता प्रति विद्यार्थी। एक exam के बाद, लगभग 2,000 वह AI job queue में थे, और हर merit job उनके पीछे बैठा।

Arithmetic scale दिखाता है। एक AI job लगभग 29 merit job जितना लंबा। दो हज़ार का मतलब लगभग 20,000 second काम, और 2 worker भर यह लगभग तीन घंटे line है। 10 to 115 मिनट का wait यह picture में fit होता है।

यह एक supermarket express lane है: bread एक full trolley के पीछे wait नहीं करना चाहिए, और हमने एक सब के लिए till बनाया।

Before और after diagram: एक shared queue जहाँ tiny merit job लगभग 2,000 ten-second AI job के पीछे wait करते, versus अलग lane जहाँ AI lane 8 per मिनट का एक shared limiter से paced है
Same job, अलग topology: दाएँ merit job कभी AI job से meet नहीं करता, और AI lane का अपना pace है।

Change

Implemented: एक अलग queue lane और इसका अपना container, सिर्फ़ AI narrative के लिए, 2 replica के साथ। Merit और position job कभी language model के लिए wait नहीं करते, और merit card लगभग एक मिनट में ready होते हैं। AI job पहले से ही queue में थे एक atomic script के साथ नए lane में move किए गए, तो हर job हर moment बिल्कुल एक जगह में था।

Fix horsepower नहीं, topology था। Generic rule एक queue per kind काम Queues और Workers at Scale में है; जो हमें surprised किया वह कैसे harmless slow job दिखता था।

फिर AI provider ने "no" कहा

नया lane merit card को protect किया। इसने AI job को faster नहीं बनाया। Lane provider का अपना rate limit, लगभग 7 to 10 calls एक मिनट, में run हुआ, और एक बार provider account आसानी से credit से बाहर चला गया।

दोनों same failure हैं: एक बाहरी service "not now" कहता है। एक retry loop जो इसे hammer करता है सिर्फ़ budget burn करता है और failed-jobs table भरता है, तो lane redesign किया गया patient बनने के लिए। Implemented और live worker पर 6 अक्टूबर से:

MechanismSettingयह क्या करता है
Shared rate limiterहर worker भर 8 per मिनटCall को provider allow करता है के अंदर रखता है
Release, fail नहींExponential backoff, 1 to 15 मिनट, jitter के साथRefused job queue में back जाता है; jitter उन्हें एक साथ return से रोकता है
Retry window12 घंटे तकJob rush के बाद लंबे समय तक trying रहता है
Budget refundMonthly AI budget counterRefused call charge नहीं होता
FallbackTemplate textविद्यार्थी template text देखता है इसी बीच

विद्यार्थी कभी error नहीं देखते। Template text वहाँ होता है जब वे page खोलते हैं, और real narrative उसे replace करता है जब इसका turn आता है।

यह real के अंदर एक दिन test किया गया। Provider account credit से बाहर चला गया, लगभग 3,800 narrative delayed job के रूप में pile हो गए, और कोई भी fail नहीं हुआ। अगले दिन शाम credit add किया गया, और पहला नया narrative मिनट के अंदर written किया गया, restart नहीं और manual replay नहीं के साथ।

Arithmetic दिखाता है क्यों window घंटे हैं: 2,000 narrative 8 एक मिनट पर लगभग चार घंटे काम है। कुछ मिनट की window ज़्यादातर को drop करता।

Limiter sharing matter करता है। दो replica जो अपने आप pace करते 4 एक मिनट पर एक तीसरा जोड़े जाने तक काम करते। एक shared counter total को honest रखता है हालांकि कितने भी worker exist करते।

Rate limit: building को नहीं, विद्यार्थी को count करो

एक rate limiter सिर्फ़ अच्छा है जितना चीज़ को count करता है। हमने यह दो बार गलत पाया, दो layer पर, और दोनों बार symptom विद्यार्थियों को exam के बीच HTTP 429 receive करना था।

कबक्या count किया गयाक्या wrong चलाFix
25 सितंबरGlobal API limiter, IP पर keyed (token-based विद्यार्थी के लिए empty default guard)एक school या mobile carrier सैकड़ों विद्यार्थी को एक NAT IP के पीछे रखता है: एक shared bucketविद्यार्थी या instructor account से key करो, IP सिर्फ़ guest के लिए (Implemented)
7 अक्टूबरRoute throttle जैसे "30 per मिनट", अब भी client IP सेBackend ने एक frontend proxy IP देखा सब के लिए: एक dashboard endpoint के 74% call को 429 मिलाPer विद्यार्थी per route count करो, guest IP से; 0 ऐसे 429 बाद में (Implemented)

हर browser API call frontend proxy के through जाता है, तो एक per-IP rule पूरी exam hall को एक visitor मानता है।

दोनों fix test के साथ आए जो fix के बिना fail करते हैं; दूसरा एक node के समय deployed था zero downtime के साथ। Planned: real विद्यार्थी IP frontend से backend में एक signed header में pass करो, तो हर per-IP rule और हर log line exact है।

Stale trusted proxy

एक छोटी risk पास में बैठी। Backend अब भी पुराने frontend server का IP trust करता था, जिसे हमने delete किया। अगर cloud कभी वह address किसी को दे, वे विद्यार्थी IP fake कर सकते। हमने इसे 7 अक्टूबर पर remove किया drain-one-node method के साथ, फिर से zero downtime के साथ। Implemented। एक trust list एक promise की list है; delete उन जिसका owner gone है।

Log जो node को outlive करते हैं

Autoscaling के साथ, machine जिसे आप inspect करना चाहते आमतौर पर gone होता है। तो log node को छोड़ना चाहिए जैसे लिखे जाते हैं। हमारे container सिर्फ़ stdout और stderr को write करते हैं, और Fluent Bit backend container log को managed OpenSearch, ELK-style stack में ship करता है: एक searchable जगह जो node के gone होने के बाद keep किया जाता है। Implemented backend node के लिए। Generic method Observability for High-Volume Systems में है।

यह capacity काम में अपने लिए pay किया। DigitalOcean की two-minute average 7 अक्टूबर backend peak को 63 requests per second पर डाला; per-minute log लगभग 69 कहते थे। Average peak को hide करते हैं, और capacity model को peak की ज़रूरत है।

एक gap है। Frontend node अभी भी अपने nginx log local disk पर रखते हैं, लगभग 440 MB एक दिन, और उन्हें lose करते हैं जब एक node remove किया जाता है। Planned: Frontend node पर भी Fluent Bit। अब तक, frontend वह layer है जहाँ scale-in evidence erase करता है।

हम क्या watch करते हैं, और हम इसके बारे में क्या करते हैं

हमारा ops console read-only है: यह sample करता है और कभी कुछ change नहीं करता। यह record करता है CPU per node और per container, pool CPU, MySQL thread running, connection और lock wait, Valkey memory, client और operation per second, और queue health। DigitalOcean का अपना monitoring request per second और response class add करता है load balancer पर।

एक exam के दौरान हम उस सब को stare नहीं करते। हमने एक digest read करते हैं एक बार एक मिनट: विद्यार्थी started और submitted, request और 5xx per node, pool CPU, worker load और failed job। कम संख्या, हर एक एक decision को tie किया।

Left-to-right loop: signal ops console से, load balancer metric और OpenSearch log, एक once-a-minute digest, फिर action जैसे drain node, roll back image और pre-scale, फिर re-measure
Signal एक digest में feed होते हैं, digest कुछ known move trigger करता है, और हर incident एक runbook rule के रूप में return करता है।

एक signal पर act करना

Loop के पास एक छोटा set move है, Part 2 में described:

  • एक node drain करो। Make /lb-health 503 return करे, और load balancer लगभग 30 second में नए request भेजना बंद करता है।
  • Image को roll back करो। हमारे पास last 3 अच्छे image rollback के लिए हैं, और हमारे पास कभी template roll किया exam घंटे के दौरान नहीं।
  • Pre-scale करो। Pool minimum को लगभग 45 मिनट पहले एक बड़े exam से raise करो। Reactive scaling सिर्फ़ safety net है, क्योंकि DigitalOcean का CPU metric 5 to 8 मिनट से real load को lag करता है।

7 अक्टूबर दोनों pool के full rollout में, एक uptime probe real page को hit किया हर 2 second: 219 probe, 0 error। यह loop को close करना है: change, watch, measure।

Scale-in की गलती

7 अक्टूबर real exam के दौरान, किसी ने backend pool minimum को lower किया। DigitalOcean एक serving node remove किया बिना drain किए, और लगभग 360 request 2 मिनट में failed।

यह एक human slip और एक platform behavior एक साथ था: autoscale scale-in drain नहीं करता। Rule अब runbook में है (Implemented runbook rule के रूप में): exam से पहले minimum raise करो, और सिर्फ़ बाद में lower करो। हमारे limit पता है: एक runbook rule depend करता एक exam के बीच किसी को इसे याद रखने पर। यह guard है, lock नहीं।

हमने क्या सीखा

  • Stateless node उतने ही safe हैं जितने shared store उनके पीछे हैं। Size Valkey connection के लिए, सिर्फ़ memory नहीं।
  • एक queue per kind काम। एक 0.34-second job कभी 10-second के पीछे wait नहीं करना चाहिए।
  • जब provider आपको limit करता है, call को pace करो, backoff करो jitter के साथ, घंटे के लिए retry करो, budget को refund करो और fallback दिखाओ।
  • Check करो कि हर rate limiter क्या count करता है। विद्यार्थी unit है, building नहीं और proxy नहीं।
  • Log को node से ship करो जैसे लिखे जाते हैं, और frontend gap को close करो इससे पहले cost incident।
  • अपने single point of failure write करो हर एक के साथ status। हमारा एक worker है, planned fix और pretending नहीं के साथ।
  • Autoscale scale-in drain नहीं करता। Exam से पहले minimum raise करो।

Part 5 में, series का अंतिम, हमने इसे एक साथ रखते: 7 अक्टूबर की real exam evening production संख्या के साथ, capacity model, monthly cost, और load और soak test जो हमें अभी भी अपने आप देना है।

About the Author

Anichur Rahaman

Continue Reading