<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[RSS Feed | Mayank Raj]]></title><description><![CDATA[Mayank Raj is a Staff Engineer on Stripe's Infrastructure team, writing about infrastructure, reliability, security, AI systems, cloud architecture, and builder communities.]]></description><link>https://mayankraj.com</link><generator>GatsbyJS</generator><lastBuildDate>Mon, 28 Sep 2026 14:45:50 GMT</lastBuildDate><item><title><![CDATA[Idempotency Keys: Why Retrying Without Them Is Financial Suicide]]></title><description><![CDATA[You're standing at an ATM at 11:47 PM and you press "Withdraw $100." The screen freezes where the loading spinner spins... and spins... and…]]></description><link>https://mayankraj.com/blog/idempotency-keys-exactly-once-payments</link><guid isPermaLink="false">https://mayankraj.com/blog/idempotency-keys-exactly-once-payments</guid><pubDate>Thu, 12 Feb 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;You&apos;re standing at an ATM at 11:47 PM and you press &quot;Withdraw $100.&quot; The screen freezes where the loading spinner spins... and spins... and spins. A minute later, the machine is silent. Did the transaction go through? Do you press the button again? Will you ever see your $100!? This is not a hypothetical. This is the Two Generals Problem, dressed in a bank lobby. And if you&apos;re building a payment API without understanding this then you&apos;re not building a financial system but you&apos;re building a slot machine where every retry pulls the lever on your customers&apos; accounts.&lt;/p&gt;
&lt;p&gt;Let me show you why idempotency keys are the only thing standing between &quot;reliable payments&quot; and &quot;oops... we accidentally charged you three times.&quot;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-network-is-your-enemy-and-it-always-lies&quot;&gt;The Network Is Your Enemy (And It Always Lies)&lt;/h2&gt;
&lt;p&gt;Here&apos;s the problem: when you make an HTTP request to charge a credit card, one of the three things can happen.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The request never reaches the server. Network jitter, DNS failure or a router hiccupped in Frankfurt. The payment never happened.&lt;/li&gt;
&lt;li&gt;The server processes the request, but it&apos;s thumps-up aka 200 back never reaches you. The money changed db-entries, the charge did go through. But your client thinks it failed.&lt;/li&gt;
&lt;li&gt;The server crashes mid-processing. The debit happened but the credit didn&apos;t. (Or vice versa, pick your nightmare.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;From your client&apos;s perspective, all three scenarios end up the same i.e. with timeout. The loading spinner never gets a signal to stop and you&apos;re left with a choice to either retry the request and risk a double-charge, or abandon it and risk the money vanishing into the ether.&lt;/p&gt;
&lt;p&gt;This is the Two Generals Problem, a classic distributed computing paradox. Two generals need to coordinate an attack between them but they can only communicate via messengers who might get captured. General A sends a messenger: &quot;Attack at noon.&quot; Did it reach? General B sends back: &quot;Acknowledged.&quot; Did that arrive? No amount of back-and-forth acknowledgments can provide 100% certainty. We see this everywhere from HTTP protocol to DB commits. In our payment system, the client and server are the Generals. The network doesn&apos;t owe them reliability of any sort.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;at-least-once-delivery&quot;&gt;At-Least-Once Delivery&lt;/h2&gt;
&lt;p&gt;Modern systems don&apos;t just give up when a request times out. They retry, and if built rightly with some sort of backoff. Everything from Apache Kafka, the HTTP client library you use and even your mobile app retries. This is called at-least-once delivery, and it&apos;s the industry standard for one simple reason: losing a message is worse than duplicating it. But in some cases, especially in financial operations, duplication is not better than loss. It&apos;s categorically worse.&lt;/p&gt;
&lt;p&gt;Let&apos;s say a customer clicks &quot;Pay $100&quot; on your checkout page. The request hits your server, charges their card, and updates the ledger, but before the server can respond with &quot;200: Success,&quot; a network glitch eats up the acknowledgment. The client thinks the request failed. So it retries, maybe multiple times. Without idempotency, your backend sees three separate requests and rightfully so it charges the card three times. The customer&apos;s bank account: $300 lighter and all of a sudden the discount coupon makes no sense.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-holy-grail-of-exactly-once-processing&quot;&gt;The Holy Grail of Exactly-Once Processing&lt;/h2&gt;
&lt;p&gt;You might be thinking: &quot;Can&apos;t we just build exactly once delivery and avoid this whole mess?&quot; No. It&apos;s mathematically impossible. The Two Generals Problem proves that you cannot guarantee delivery over an unreliable network. What is achievable, and what you must build is exactly-once processing. Here&apos;s the distinction:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Exactly-once delivery would mean the message arrives at the recipient exactly one time. In theory the network guarantees this, in reality it can&apos;t.&lt;/li&gt;
&lt;li&gt;Exactly-once processing means the side effects of a message occur exactly once, even if the message itself is delivered multiple times.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the entire game. You accept the fact that the network will deliver duplicates at-least-once, but you make your server smart enough to recognize: &quot;Looks like I&apos;ve already processed this request. Here&apos;s the result I cached last time.&quot;&lt;/p&gt;
&lt;p&gt;That intelligence is called idempotency. And the key to unlocking it is... well.. the idempotency key.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-idempotency-key-a-deterministic-anchor-in-a-chaotic-sea&quot;&gt;The Idempotency Key: A Deterministic Anchor in a Chaotic Sea&lt;/h2&gt;
&lt;p&gt;An idempotency key is a unique identifier the client generates before making a request. It&apos;s typically a UUID v4, attached as an HTTP header:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-http&quot;&gt;POST /api/payments
Idempotency-Key: 7f3e5c1a-9b2d-4f6e-8ae4-3f4e5f6a7b8c
Content-Type: application/json

{
  &quot;amount&quot;: 10000,
  &quot;currency&quot;: &quot;USD&quot;,
  &quot;user_id&quot;: &quot;cust_12345&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;When the server receives this request, it follows a specific protocol:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Check for existence. Query the idempotency store (database or cache): &quot;Have I seen this key before?&quot;&lt;/li&gt;
&lt;li&gt;If yes, return the cached result. We don&apos;t re-execute the charge or touch the database.&lt;/li&gt;
&lt;li&gt;If no, process the request. Execute the business logic, save the response in the idempotency store, and then return it to the client.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, when the client retries because it didn&apos;t receive the response (scenario #2 from earlier), the server recognizes the key. It skips the charge logic entirely. It returns the cached success response. The customer is charged once. The ledger is consistent. The idempotency key converts an unreliable at-least-once delivery into deterministic exactly-once processing.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;why-the-client-generates-the-key-not-the-server&quot;&gt;Why the Client Generates the Key (Not the Server)&lt;/h2&gt;
&lt;p&gt;You might wonder: &quot;Why doesn&apos;t the server just generate the key and return it to the client?&quot;&lt;/p&gt;
&lt;p&gt;Because that would reintroduce the problem at the key-exchange layer. If the server generates the key and sends it to the client, and that response gets lost, the client is blind to the key. It&apos;s as if it never existed. We&apos;ve moved the problem around and not solved it. By having the client generate the key before the first attempt, the key becomes a deterministic anchor that survives network failures, app crashes, and even client-side restarts. The client can persist the key to local storage or a cookie. So if the app or webpage crashes mid-request, the next session can reload the key and retry safely.&lt;/p&gt;
&lt;p&gt;Stripe does this. Airbnb does this. Every competent payment API does this. (And if your API doesn&apos;t, please stop reading and go implement it. I&apos;ll wait.)&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;request-fingerprinting-preventing-key-reuse-exploits&quot;&gt;Request Fingerprinting: Preventing Key Reuse Exploits&lt;/h2&gt;
&lt;p&gt;Here&apos;s a nasty edge case: what if a malicious (or buggy, but mostly malicious) client reuses an idempotency key from a $10 payment for a $1,000 payment? If your server just blindly returns the cached $10 success response for the $1,000 request, you just handed a &quot;$990 off&quot; coupon, and without any coupon code. The fix is to store a cryptographic hash of the request payload alongside the idempotency key.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import hashlib
import json

def compute_request_hash(method, uri, body):
    # Combine everything that makes this request unique
    payload = f&quot;{method}:{uri}:{json.dumps(body, sort_keys=True)}&quot;
    return hashlib.sha256(payload.encode()).hexdigest()  # Deterministic fingerprint
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;When a request with a known idempotency key arrives, compare the stored hash to the incoming request&apos;s hash:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If they match: Return the cached response.&lt;/li&gt;
&lt;li&gt;If they don&apos;t match: Return &lt;code&gt;HTTP 409 Conflict&lt;/code&gt; and force the client to use a new key.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This ensures that an idempotency key is uniquely and irrevocably bound to a specific set of parameters. You can retry the same charge safely but you cannot reuse the key for a different charge. This also means anything that can change with time, say &lt;code&gt;current_timestamp&lt;/code&gt; cannot be part of this hash.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;http-status-codes-when-to-cache-when-to-retry&quot;&gt;HTTP Status Codes: When to Cache, When to Retry&lt;/h2&gt;
&lt;p&gt;Not all errors should be cached under an idempotency key. Here&apos;s a common policy:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status Code&lt;/th&gt;
&lt;th&gt;Should Cache?&lt;/th&gt;
&lt;th&gt;Client Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2xx Success&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Done. Transaction succeeded.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;400 Bad Request&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Invalid parameters. Fix data, use &lt;em&gt;new&lt;/em&gt; key.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;409 Conflict&lt;/td&gt;
&lt;td&gt;❌ No&lt;/td&gt;
&lt;td&gt;Idempotency key reused incorrectly. Use new key.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;429 Rate Limit&lt;/td&gt;
&lt;td&gt;❌ No&lt;/td&gt;
&lt;td&gt;Retry with &lt;em&gt;same&lt;/em&gt; key after backoff.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5xx Server Error&lt;/td&gt;
&lt;td&gt;Usually ❌ No&lt;/td&gt;
&lt;td&gt;Transient failure. Retry with same key unless your API explicitly persists failures.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Why cache 400s? Because if the client sent invalid data and retrying with the same key won&apos;t magically fix it. Caching the error prevents pointless retries and clarifies: &quot;This request is fundamentally broken.&quot; So why not cache 5xx by default? Because the server might have crashed mid-transaction or a microservice was unavailable. A retry might succeed once the database recovers. That said, some APIs intentionally persist the first 5xx result to preserve strict idempotency semantics, so this part is a policy choice, not a law of physics.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;key-expiration&quot;&gt;Key Expiration&lt;/h2&gt;
&lt;p&gt;Storing idempotency keys forever is impractical and by design not required in the majority of cases. Your database would grow unbounded, and query performance would degrade. Most systems implement a time-to-live (TTL) for such keys. Say the key expires after 24 hours. This window is sufficient for any reasonable retry scenario. If a client retries after 24 hours then the key is gone, and the request is treated as new. Which is probably fine and if your retry loop is still running a day later, you have bigger problems to solve.&lt;/p&gt;
&lt;p&gt;For high-stakes operations (like inter-bank transfers), you might extend the TTL to 7 days or store keys indefinitely in cold storage. The trade-off is between storage cost and the risk of re-executing a very old request.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;case-study-airbnbs-orpheus-framework&quot;&gt;Case Study: Airbnb&apos;s Orpheus Framework&lt;/h2&gt;
&lt;p&gt;Airbnb&apos;s transition to microservices required a centralized idempotency solution. They built &lt;a href=&quot;https://medium.com/airbnb-engineering/avoiding-double-payments-in-a-distributed-payments-system-2981f6b070bb&quot;&gt;Orpheus&lt;/a&gt;, a library that models every API request in three phases:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Pre-RPC Phase: Record the intent to perform an action. This happens in a database transaction alongside local state changes.&lt;/li&gt;
&lt;li&gt;RPC Phase: Make the external call (e.g., to Stripe or Braintree). This is inherently non-atomic because the external service is a separate system.&lt;/li&gt;
&lt;li&gt;Post-RPC Phase: Record the result of the external call. Update the idempotency record with the final response.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If a crash occurs during the Pre-RPC phase then the transaction rolls back and no side effects occurred. The client can retry safely. If a crash occurs during or after the RPC phase, the idempotency key allows recovery. On retry, the library checks the database and if it finds a recorded RPC result, it skips the external call and proceeds directly to the Post-RPC phase. This three-phase model ensures that guests are never double charged and hosts are always paid correctly, even when the network is actively conspiring against such holidays.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-secret-idempotency-is-a-pinky-promise-not-a-feature&quot;&gt;The Secret: Idempotency Is a Pinky Promise, Not a Feature&lt;/h2&gt;
&lt;p&gt;Idempotency keys aren&apos;t just a technical safeguard but they&apos;re a promise to your users that you understand the network is unreliable, and you&apos;ve designed your system to be resilient anyway. When a customer clicks &quot;Pay,&quot; they&apos;re trusting you with their money. They&apos;re trusting that your system is deterministic, even when the network is chaotic. They&apos;re trusting that you&apos;ve done your homework. That trust is the foundation of every financial platform - Stripe, PayPal, Airbnb, Shopify, Amazon... you name it. These all use idempotency keys because they understand that in a distributed system, retrying without idempotency is not a technical risk but it&apos;s a betrayal of user trust.&lt;/p&gt;
&lt;p&gt;That&apos;s not just good engineering. That&apos;s the only engineering that matters.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Why Your Cloud Bill Keeps Growing (And Why That Might Be Fine): The Jevons Paradox in Cloud Computing]]></title><description><![CDATA[You did it! You finally did it!! Months of planning, countless nights, POCs but your team migrated to serverless. You implemented auto…]]></description><link>https://mayankraj.com/blog/jevons-paradox-cloud-efficiency-trap</link><guid isPermaLink="false">https://mayankraj.com/blog/jevons-paradox-cloud-efficiency-trap</guid><pubDate>Wed, 21 Jan 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;You did it! You finally did it!! Months of planning, countless nights, POCs but your team migrated to serverless. You implemented auto-scaling, and optimized every last Lambda function. The unit costs deserve chef&apos;s kiss with each transaction now costing 42% less. You present the metrics to leadership, expecting applause and a hefty bonus.&lt;/p&gt;
&lt;p&gt;Then the monthly bill arrives. It&apos;s 30%... higher than before, not lower - but higher. Let&apos;s look at why that&apos;s a even more beautiful metric.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Here&apos;s the thing about highways: adding more lanes doesn&apos;t reduce traffic congestion. You instead end up with even more traffic. Civil engineers call this phenomenon induced demand. The moment you make driving more convenient, more people drive. Those who were escaping rush hour traffic with train, now switch to cars. Suburban sprawl accelerates because the commute is now manageable. The new lanes fill up, and you&apos;re right back where you started, except now you&apos;re maintaining twice as much asphalt.&lt;/p&gt;
&lt;p&gt;Highways are the cloud abstractions you built. Cars are the services running on them. Making it easy to run services increases the throughput, and with it the cost. Your cloud infrastructure works exactly the same way. Maybe less public commute is a bad thing for the city, but for you ? A higher bill could be a healthy indicator.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-coal-question-160-years-later&quot;&gt;The Coal Question (160 Years Later)&lt;/h2&gt;
&lt;p&gt;In 1865, William Stanley Jevons published &lt;a href=&quot;https://en.wikipedia.org/wiki/The_Coal_Question&quot;&gt;The Coal Question&lt;/a&gt;, a book that should be required reading for every engineering leader with a straightforward view on optimizing cloud spend. Jevons observed something counterintuitive: when James Watt&apos;s steam engine made coal power more efficient than Thomas&apos;s design, England&apos;s total coal consumption didn&apos;t decrease. It exploded. The efficiency didn&apos;t lead to conservation but it made coal so cost-effective and versatile that industries found thousands of new applications for it. Everything from blast furnaces to shipping to textile mills. The resource became cheaper per unit of work, so the economy consumed vastly more work.&lt;/p&gt;
&lt;p&gt;This is now called the Jevons Paradox: improvements in efficiency lead to increases in total consumption. Today you&apos;re not burning coal but you&apos;re burning compute cycles. But the paradox? Still alive. Still expensive. Still smiling at your face.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-efficiency-trap&quot;&gt;The Efficiency Trap&lt;/h2&gt;
&lt;p&gt;Let me show you how this plays out in practice. A financial services company (let&apos;s call them &quot;FinCorp&quot;) migrated from EC2 instances to a fully serverless architecture on AWS Lambda. The results were impressive:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Per-transaction cost: Down 42%&lt;/li&gt;
&lt;li&gt;Infrastructure team time: Reduced by 60%&lt;/li&gt;
&lt;li&gt;Response latency: Improved to P99 &amp;#x3C; 100ms&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The CFO was thrilled. &quot;If we&apos;re processing the same transactions at 42% lower unit cost, at least our compute bill should drop by nearly half, right?&quot; Wrong. Here&apos;s what actually happened:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;The Direct Rebound: Now that the transactions are cheaper, the product team launched a real time fraud detection service that was previously &quot;too expensive&quot; to run. Each transaction now triggers three additional Lambda invocations for risk scoring.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The Experimental Sprawl: Developers are no longer constrained by &quot;the fixed pie&quot; of EC2 capacity. They spun up 47 new microservices to test personalization features, A/B experiments, and ML model variations.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The Architectural Expansion: With auto-scaling removing the fear of capacity planning, the team added real-time dashboards. WebSocket connections for live updates, and a new mobile API that polls every 30 seconds are now in production.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;While FinCorp&apos;s cloud bill instead increased by 100%, they&apos;re serving 400% more customers. On top of it three new revenue streams were launched. The CFO, once skeptical, now sees cloud spend as a growth indicator, not a cost center.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-economics-why-efficiency-breeds-consumption&quot;&gt;The Economics: Why Efficiency Breeds Consumption&lt;/h2&gt;
&lt;p&gt;To understand why this keeps happening time and again, we need to talk about price elasticity of demand. This is the relationship between the price of something and how much of it people consume. The formula is straightforward:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Elasticity = |Δ Quantity / Quantity| / |Δ Price / Price|
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;When demand is elastic (elasticity &gt; 1), a drop in price causes a more than proportional increase in consumption. If your serverless migration cuts costs by 50%, but your usage increases by 150%, you&apos;ve just experienced the Jevons Paradox.&lt;/p&gt;
&lt;p&gt;Here&apos;s where it gets interesting: traditional IT infrastructure has had low elasticity. With physical servers, a price drop doesn&apos;t immediately allow you to buy more. You&apos;re constrained by procurement lead times, data center space, and the administrative overhead of racking hardware.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Infrastructure Model&lt;/th&gt;
&lt;th&gt;Lead Time&lt;/th&gt;
&lt;th&gt;Financial Model&lt;/th&gt;
&lt;th&gt;Elasticity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Physical Servers&lt;/td&gt;
&lt;td&gt;Weeks to Months&lt;/td&gt;
&lt;td&gt;CapEx&lt;/td&gt;
&lt;td&gt;Highly Inelastic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Colo/Rental&lt;/td&gt;
&lt;td&gt;Days&lt;/td&gt;
&lt;td&gt;OpEx&lt;/td&gt;
&lt;td&gt;Modestly Inelastic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Public Cloud&lt;/td&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;td&gt;Pay-as-you-go&lt;/td&gt;
&lt;td&gt;Elastic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The public cloud removed nearly all barrier to consuming more compute. Procurement, capacity planning is gone. There&apos;s no upfront capital. Just spin up another Lambda function and the AWS bill adjusts itself. This is why 83% of organizations &lt;a href=&quot;https://www.azul.com/newsroom/azul-report-finds-83-of-cios-are-spending-more-on-their-cloud-infrastructure-and-applications-than-anticipated/&quot;&gt;report&lt;/a&gt; spending more on cloud services than anticipated, with an average overspend of 30%. The infrastructure isn&apos;t failing but it&apos;s working exactly as designed. You just unlocked infinite demand and means to fulfil it.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;architectural-sprawl-when-efficiency-enables-chaos&quot;&gt;Architectural Sprawl: When Efficiency Enables Chaos&lt;/h2&gt;
&lt;p&gt;The Jevons Paradox doesn&apos;t just affect infrastructure but it also infects your architecture. In the old world of monolithic applications running on a somewhat fixed EC2 capacity, developers were constrained. You had 64GB of RAM and 16 vCPUs. If your code was inefficient, it impacted everyone else on the box. Scarcity bred discipline. And discipline called for sane choices upfront.&lt;/p&gt;
&lt;p&gt;Serverless removed the constraint. Now, adding a new feature doesn&apos;t require a capacity planning meeting. It&apos;s just another Lambda. The marginal cost of experimentation approaches zero, so the organization experiments... a lot. This creates what we can call &quot;architectural sprawl&quot; - a silent but impactful explosion of microservices. Now there&apos;s one additional deployment pipeline, monitoring dashboard, and dependency graph.&lt;/p&gt;
&lt;p&gt;Let&apos;s say you decompose your Django monolith into 72 microservices. Each service makes sense in isolation with clean boundaries, independent deployment, small team ownership. But the system as a whole became a distributed monolith with extra steps and with extra Lambda invocations.&lt;/p&gt;
&lt;p&gt;A single user login now triggers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;1 auth service call&lt;/li&gt;
&lt;li&gt;3 profile enrichment calls&lt;/li&gt;
&lt;li&gt;2 analytics events&lt;/li&gt;
&lt;li&gt;4 feature flag checks&lt;/li&gt;
&lt;li&gt;1 A/B experiment assignment&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That&apos;s ~11 Lambda invocations for what used to be a single database query. The cost per login, latency, overall load, they all went up. The blast radius is now bigger. But here&apos;s the uncomfortable truth: you&apos;re now shipping features 10x faster than before. The architectural sprawl is a feature, not a bug. Your business traded operational simplicity for velocity, and velocity won.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-counter-argument-maybe-the-paradox-has-limits&quot;&gt;The Counter-Argument: Maybe the Paradox Has Limits&lt;/h2&gt;
&lt;p&gt;Before you despair, let&apos;s acknowledge the skeptics. Not everyone believes the Jevons Paradox applies unboundedly, especially now with the AI hype and cloud computing.&lt;/p&gt;
&lt;p&gt;Here are the arguments for why consumption might eventually plateau at some point:&lt;/p&gt;
&lt;h3 id=&quot;1-market-saturation&quot;&gt;1. Market Saturation&lt;/h3&gt;
&lt;p&gt;LED lighting initially exhibited a rebound effect wherein people left lights on longer because they were cheap to run. But the market eventually saturated. You can only leave the lights on 24 hours a day. Is there a ceiling to how much AI generated slop content or real-time analytics a business can actually use?&lt;/p&gt;
&lt;h3 id=&quot;2-diminishing-returns&quot;&gt;2. Diminishing Returns&lt;/h3&gt;
&lt;p&gt;At some point, adding more microservices or running more ML experiments stops generating proportional business value. The ROI curve flattens and the capital allocation shifts to other investments.&lt;/p&gt;
&lt;h3 id=&quot;3-physical-constraints&quot;&gt;3. Physical Constraints&lt;/h3&gt;
&lt;p&gt;Unlike 19th-century coal, AI compute requires specialized silicon (GPUs, TPUs, and whatever Groq has) and massive CapEx for data center construction. Supply chains have limits and so does Corporate budgets.&lt;/p&gt;
&lt;h3 id=&quot;4-regulatory-ceilings&quot;&gt;4. Regulatory Ceilings&lt;/h3&gt;
&lt;p&gt;Governments are waking up to the environmental cost of AI. We cannot keep milking the cow. If carbon taxes or energy caps become common, the paradox could hit a hard regulatory wall.&lt;/p&gt;
&lt;p&gt;These are real constraints. The Jevons Paradox isn&apos;t a law of physics but it&apos;s an economic tendency. Whether it holds in the long term depends on how markets, technologies, and policies evolve. But for the next decade? I&apos;ll still bet on the paradox.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-takeaway-optimize-for-value-not-cost&quot;&gt;The Takeaway: Optimize for Value, Not Cost&lt;/h2&gt;
&lt;p&gt;Here&apos;s what all of us as leaders need to internalize:&lt;/p&gt;
&lt;p&gt;Making your systems more efficient will not reduce your total cloud spending. It will make compute cheaper, which in turn will unlock new use cases, which will drive higher total consumption. This is not a failure of cost optimization but it&apos;s the natural outcome of lowering barriers to innovation.&lt;/p&gt;
&lt;p&gt;Your job should not be just to minimize the AWS bill. Your job should be to ensure that every dollar spent generates maximum business value. That&apos;s a metric worth spending time on. This requires a shift in mindset:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Stop measuring absolute spend and start measuring cost per unit of value delivered (cost per transaction, cost per customer, cost per inference).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Embrace FinOps as a culture, not a project. Give engineering teams real-time visibility into their unit economics and empower them to make trade-offs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Govern architectural sprawl. Efficiency enables innovation, but unchecked sprawl eventually leads to diminishing returns. Enforce principles around service granularity and dependency management.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The enterprises that thrive in the cloud era won&apos;t be the ones that spend the least. They&apos;ll be the ones that generate the most value from every Joule of energy and every millisecond of compute they consume.&lt;/p&gt;
&lt;p&gt;The Jevons Paradox teaches us that efficiency isn&apos;t about using less. It&apos;s about doing more with what you have. The highway lanes will always fill up. The real question is: where are all those new drivers going?&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Stateless Lie: How AWS Lambda Finally Learned to Remember]]></title><description><![CDATA[Remember making long-distance phone calls in the 1990s? You'd have to plan the conversation like it's a military operation, rehearse your…]]></description><link>https://mayankraj.com/blog/lambda-durable-functions-stateless-lie</link><guid isPermaLink="false">https://mayankraj.com/blog/lambda-durable-functions-stateless-lie</guid><pubDate>Fri, 26 Dec 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Remember making long-distance phone calls in the 1990s? You&apos;d have to plan the conversation like it&apos;s a military operation, rehearse your key points, and watch the clock with existential dread. Every minute costed real money. Being disconnected midway meant that you&apos;d have to call back, re-explain the context, and hope you remembered where you left off. The anxiety wasn&apos;t about the technology failing but the architectural constraint that made natural conversations feel like a series of frantic sprints against time.&lt;/p&gt;
&lt;p&gt;I remember discussing this exact constraint with a Principal Engineer from AWS during one of my interviews. We were talking about Lambda&apos;s 15-minute limit, and they were very clear: &quot;We&apos;re smart enough to remove that limit, but there has to be a good reason behind it.&quot; Then they paused and added, &quot;I&apos;m sure needs will evolve ...and so will Lambda.&quot; At the time, I nodded like I understood the depth of what they meant. (well I didn&apos;t.)&lt;/p&gt;
&lt;p&gt;For well over a decade, AWS Lambda has lived in that same 1990s phone call universe. Every function had a hard limit of 15 minutes to say what it needed to say, and if your business logic took longer ...well, you&apos;d split it. You built external orchestrators, wrote &quot;Lambda calling Lambda&quot; chains. We had the whole of AWS Step Functions come in. We learned to express the application logic in Amazon States Language. Because nothing says &quot;developer productivity&quot; like maintaining separate infrastructure-as-code files for what should&apos;ve been a simple &lt;code&gt;while&lt;/code&gt; loop. We termed this &quot;stateless computing&quot; and treated it like some sort of immutable law of physics rather. In reality it was an architectural guardrail we&apos;d outgrown.&lt;/p&gt;
&lt;p&gt;At AWS re:Invent 2025, that guardrail came down.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-15-minute-wall-wasnt-technical-it-was-philosophical&quot;&gt;The 15-Minute Wall Wasn&apos;t Technical ...It Was Philosophical&lt;/h2&gt;
&lt;p&gt;When Lambda launched in 2014, the 15-minute execution limit was presented as a feature, not a bug. We all brought it. Short-lived, atomic, idempotent functions were the doctrine. The reasoning made sense at the time: prevent runaway processes all while recycling the compute resources efficiently, and force developers to think in stateless, event-driven terms.&lt;/p&gt;
&lt;p&gt;This was the era of microservices ascendant, where decoupling and independence were the answers to every architectural question. I caved in all well. The term &quot;Serverless&quot; was in my LinkedIn profile for years.&lt;/p&gt;
&lt;p&gt;But real-world business processes are &lt;strong&gt;inconveniently continuous&lt;/strong&gt;. Rarely does order fulfillment doesn&apos;t fit in 15 minutes, or that CSV processing. Multi-stage approval workflows laugh at arbitrary time limits. Chaining multiple Large Language Model calls together where each model might take 2-3 minutes and fail intermittently due to rate limiting. You&apos;d checkpoint progress in DynamoDB, overengineer SQS queues for continuation signals, and write enough error-handling logic to make a Haskell programmer nod approvingly. (Okay, I take that back.)&lt;/p&gt;
&lt;p&gt;The cognitive overhead was real. We called it the &quot;context switching tax&quot;. While your business logic may have lived in code, but the flow of that logic lived in Step Functions configuration. Want to add a simple retry with exponential backoff? That&apos;s not a &lt;code&gt;try/catch&lt;/code&gt; block anymore; that&apos;s a state machine definition update. The code and the coordination were architecturally divorced, and developers paid alimony in YAML ...or JSON.&lt;/p&gt;
&lt;p&gt;AWS Lambda Durable Functions changes the game from ground up. A single function can now run for up to one year!!! All of this while preserving state across interruptions. The &quot;pause&quot; and &quot;resume&quot; is handled directly in the runtime. This isn&apos;t just a longer timeout but instead it&apos;s a fundamentally different programming model.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;how-do-you-pause-a-function-for-six-months-without-burning-money&quot;&gt;How Do You Pause a Function for Six Months Without Burning Money?&lt;/h2&gt;
&lt;p&gt;AWS came through with an elegant solution here. The &lt;strong&gt;checkpoint and replay mechanism&lt;/strong&gt;. Unlike a traditional Lambda that executes linearly and discards its memory state when done, a durable function represents a complete lifecycle that can survive multiple physical invocations.&lt;/p&gt;
&lt;p&gt;Here&apos;s the trick: When your function hits a &quot;durable operation&quot;, say, a &lt;code&gt;context.wait(Duration.days(180))&lt;/code&gt; then the Lambda doesn&apos;t keep a warm compute instance idling. Instead, it:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Captures a checkpoint&lt;/strong&gt; of the execution state (which variables hold what values, where in the code you are)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Terminates the compute environment&lt;/strong&gt; (you stop paying for idle time)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Schedules a wake up event&lt;/strong&gt; for 180 days later&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Re-invokes the function&lt;/strong&gt; when the wait expires&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now here&apos;s where it gets super interesting. When the function wakes back up, Lambda doesn&apos;t magically restore memory. Instead it &lt;strong&gt;replays the entire function from the beginning&lt;/strong&gt;!! The SDK intercepts calls to durable operations like &lt;code&gt;step()&lt;/code&gt; or &lt;code&gt;wait()&lt;/code&gt; and lies to them. Instead of re-executing, it returns the cached results from the checkpoint store. The code speeds through its own history like a student cramming before an exam, catching up to where it was suspended, then continues with fresh compute.&lt;/p&gt;
&lt;p&gt;Let&apos;s see what this looks like in practice:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;import { DurableClient } from &quot;@aws/durable-execution-sdk&quot;;

export const handler = DurableClient.handle(async (context) =&gt; {
  // Gets replayed every time, so better be quick about it
  const orderId = context.input.orderId;

  // Step 1: Call payment service (the SDK remembers this if we crash)
  const payment = await context.step(&quot;process-payment&quot;, async () =&gt; {
    // Only THIS step retries on failure, not the whole saga
    return await chargeCustomer(orderId);
  });

  // Step 2: Now we hibernate. No compute running. Just... waiting.
  await context.waitForCallback(&quot;warehouse-check&quot;);

  // Step 3: We wake up here when the warehouse finally responds
  await context.step(&quot;ship-order&quot;, async () =&gt; {
    return await scheduleShipment(orderId);
  });

  return { status: &quot;completed&quot;, orderId };
});
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That &lt;code&gt;waitForCallback()&lt;/code&gt; is doing something magical. It generates a unique callback token, hibernates the function, and waits potentially for hours or days. This wait ends when an external system invokes a specific API with that token. When the warehouse system finally confirms availability, Lambda wakes the function up, replays it past the first two operations, and continues. No polling. No SQS queue management. No separate state machine tracking which step you&apos;re on.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-determinism-tax-your-code-runs-multiple-times-so-it-better-get-the-same-answer&quot;&gt;The Determinism Tax: Your Code Runs Multiple Times, So It Better Get The Same Answer&lt;/h2&gt;
&lt;p&gt;This replay architecture introduces a constraint that will make some developers uncomfortable: &lt;strong&gt;your code must be deterministic&lt;/strong&gt;. The function is re-executed multiple times for the same logical workflow, any logic outside of a durable &lt;code&gt;step()&lt;/code&gt; must produce exactly identical results on every replay.&lt;/p&gt;
&lt;p&gt;Consider the failure mode: You call &lt;code&gt;Math.random()&lt;/code&gt; outside a step. On the first execution it generates &lt;code&gt;&quot;abc123&quot;&lt;/code&gt;. The function completes the step and checkpoints. Later when it replays, &lt;code&gt;Math.random()&lt;/code&gt; generates &lt;code&gt;&quot;xyz789&quot;&lt;/code&gt;. The SDK is now panicking trying to reconcile checkpointed history with different data. The execution fails with a &quot;replay divergence&quot; error, and you&apos;re debugging distributed systems behavior at 2 AM. (Your on-call rotation thanks you.)&lt;/p&gt;
&lt;p&gt;The fix: &lt;strong&gt;move non-deterministic operations inside durable steps&lt;/strong&gt;. Wrap &lt;code&gt;Math.random()&lt;/code&gt;, &lt;code&gt;Date.now()&lt;/code&gt;, external API calls or anything that might change in &lt;code&gt;context.step()&lt;/code&gt;. The result is checkpointed once, then cached for all subsequent replays.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Deterministic Hazard&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Why It Breaks&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Solution&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Math.random()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Different value on every replay&lt;/td&gt;
&lt;td&gt;Wrap in &lt;code&gt;context.step()&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Date.now()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Time keeps moving forward&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;context.timestamp&lt;/code&gt; or wrap in a step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Global variables&lt;/td&gt;
&lt;td&gt;Might change between replays&lt;/td&gt;
&lt;td&gt;Pass state through function arguments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;External API calls&lt;/td&gt;
&lt;td&gt;Network is a lie&lt;/td&gt;
&lt;td&gt;Always wrap in &lt;code&gt;context.step()&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Iterating over &lt;code&gt;Map&lt;/code&gt; or &lt;code&gt;Set&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Iteration order can vary by runtime&lt;/td&gt;
&lt;td&gt;Use arrays or ensure stable ordering&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This determinism requirement is the new tax of admission. You&apos;re getting a function that can pause for months and in exchange you&apos;re writing code with re-entrancy discipline. For teams used to stateless Lambdas where every invocation is a fresh slate, this is a mindset shift.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-economic-pivot-when-durable-functions-are-cheaper-than-step-functions&quot;&gt;The Economic Pivot: When Durable Functions Are Cheaper Than Step Functions&lt;/h2&gt;
&lt;p&gt;Let&apos;s talk money, because architecture decisions that ignore cost are just expensive blog posts.&lt;/p&gt;
&lt;p&gt;AWS Step Functions bill on state transitions at &lt;strong&gt;$25 per million transitions&lt;/strong&gt;. Lambda Durable Functions use a different model: &lt;strong&gt;$8 per million durable operations&lt;/strong&gt; plus standard Lambda compute cost. While you&apos;re waiting, you pay nothing for compute.&lt;/p&gt;
&lt;p&gt;For a concrete example: Processing 10,000 items in a loop costs ~$0.75 with Step Functions (3 state transitions per item) versus ~$0.10 with Durable Functions (compute + durable operations). For high-frequency loops with minimal CPU overhead, Durable Functions can be &lt;strong&gt;~8x cheaper&lt;/strong&gt;. Step Functions charges you for every state transition while Durable Functions charges you for the durable checkpoints and actual compute consumed.&lt;/p&gt;
&lt;p&gt;(Of course, if your workflow is 3 steps with a 6-month wait between each, both cost almost nothing. At that &quot;scale&quot;, you&apos;re optimizing for developer sanity and not pennies.)&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-ai-catalyst-why-this-matters-now&quot;&gt;The AI Catalyst: Why This Matters Now&lt;/h2&gt;
&lt;p&gt;Timing of Durable Functions isn&apos;t just coincidental. The primary driver for this &quot;stateless-to-stateful&quot; shift is &lt;strong&gt;Generative AI&lt;/strong&gt; and agentic workflows. AI-powered applications fundamentally break the 15-minute model in ways that traditional business logic never did.&lt;/p&gt;
&lt;h3 id=&quot;llm-chaining-and-resilience&quot;&gt;LLM Chaining and Resilience&lt;/h3&gt;
&lt;p&gt;Chaining multiple Large Language Model calls together often takes several minutes per call and is subject to intermittent API failures or rate limiting. In a stateless Lambda, a failure at the third step of a five-step chain means restarting the entire process. This means wasting both time and token costs (which for frontier models are not trivial).&lt;/p&gt;
&lt;p&gt;With Durable Functions, we checkpoint after each model call:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;async def my_billion_dollar_funded_ai_research_agent(context: DurableContext):
    # Step 1: Generate questions (2-3 min, $0.15 in tokens we&apos;d rather not waste)
    questions = await context.step(&quot;generate-questions&quot;,
        lambda: call_claude_opus(prompt=&quot;Generate 5 research questions&quot;))

    # Step 2: Fan out to search APIs (they WILL rate limit us)
    results = await context.parallel([
        context.step(f&quot;search-{i}&quot;, lambda q=q: search_api(q))
        for i, q in enumerate(questions)
    ])

    # Step 3: Synthesize (another 3-5 min, another $0.20 we don&apos;t want to repeat)
    synthesis = await context.step(&quot;synthesize&quot;,
        lambda: call_claude_opus(prompt=f&quot;Synthesize: {results}&quot;))

    return synthesis
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If the synthesis step fails due to rate limiting, only that step retries. We don&apos;t re-run the expensive question generation or the search API calls. The previous results are cached in the checkpoint. For AI workflows where individual LLM calls cost dollars in API fees, this resilience model is a direct cost saver.&lt;/p&gt;
&lt;h3 id=&quot;human-in-the-loop-ai&quot;&gt;Human-in-the-Loop AI&lt;/h3&gt;
&lt;p&gt;Agentic AI often requires a &quot;human-in-the-loop&quot; to verify a high-stakes decision or provide domain expertise. Historically, this required complex external state machines to pause the workflow, send a notification, and resume hours or days later when the human responded.&lt;/p&gt;
&lt;p&gt;With Durable Functions, the AI agent simply waits:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;export const handler = DurableClient.handle(async (context) =&gt; {
  const analysis = await context.step(&quot;analyze-contract&quot;, async () =&gt; {
    // AI analyzes a legal contract, takes 5 minutes
    return await llm.analyze(context.input.contractText);
  });

  if (analysis.confidence &amp;#x3C; 0.85) {
    // Send notification to lawyer, wait for their review
    const reviewUrl = context.getCallbackUrl(&quot;legal-review&quot;);
    await sendEmail(lawyer, `Review needed: ${reviewUrl}`);

    // Function hibernates here (could be hours or days)
    const humanReview = await context.waitForCallback(&quot;legal-review&quot;);

    // Resume with human feedback
    return await context.step(&quot;final-decision&quot;, async () =&gt; {
      return await llm.finalizeDecision(analysis, humanReview);
    });
  }

  return analysis;
});
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The lawyer clicks the link, provides feedback via a web form and the function resumes exactly where it left off. No Step Functions state machine tracking &quot;which branch are we in.&quot; No DynamoDB table storing intermediate state. Just - plain - simple - code - that - works.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-coin-flip-when-durable-functions-are-the-wrong-choice&quot;&gt;The Coin Flip: When Durable Functions Are The Wrong Choice&lt;/h2&gt;
&lt;p&gt;This wouldn&apos;t be a balanced architectural analysis without acknowledging where Durable Functions absolutely &lt;strong&gt;shouldn&apos;t&lt;/strong&gt; be used.&lt;/p&gt;
&lt;p&gt;Please &lt;strong&gt;don&apos;t over-engineer them for simple infrastructure orchestration.&lt;/strong&gt; If your workflow is &quot;write to DynamoDB, then invoke Lambda, then send SNS notification,&quot; Step Function&apos;s direct AWS service integrations are cleaner. Step Functions can do this entirely in a state machine definition and are built for this - now more so than ever.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Don&apos;t use them when non-technical stakeholders need to debug failures.&lt;/strong&gt; The graphical Step Functions execution graph is irreplaceable for on-call engineers or support teams who didn&apos;t write the code. Durable Functions require logging, reading logs and understanding code structure. (Good luck explaining CloudWatch Insights queries to your VP of Operations.) If your workflows span multiple business units and require &quot;management-friendly&quot; observability, Step Function with their visual language justifies the cost premium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Don&apos;t use them if your team isn&apos;t ready for the determinism discipline.&lt;/strong&gt; The replay mechanism is elegant but can be unforgiving. If your developers aren&apos;t comfortable with re-entrancy, checkpointing semantics, and the constraints of deterministic code, the onboarding curve is steep. Organizations should pilot Durable Functions with a small tasks before migrating critical workflows.&lt;/p&gt;
&lt;p&gt;Finally - &lt;strong&gt;Don&apos;t use them for workflows that are already well-served by Step Functions.&lt;/strong&gt; If you have mature, battle-tested state machines that work - then let it work! The migration cost (rewriting ASL in code, testing replay behavior, training teams) might not justify the benefits. Durable Functions shine for &lt;strong&gt;new&lt;/strong&gt; workflows or workflows where Step Functions was always an awkward fit.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;well-serverless-has-grown-up&quot;&gt;Well ...Serverless Has Grown Up&lt;/h2&gt;
&lt;p&gt;The &quot;stateless serverless&quot; paradigm was never about technical necessity. It was an architectural training wheel. For a decade, we&apos;ve been forced to fragment naturally continuous logic across multiple services all while accepting the context-switching tax. AWS Lambda Durable Functions represent the maturity of serverless computing. The programming model is now closer to how we think about business processes: as continuous flows with pauses, retries, and branches, not as disconnected atomic steps.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The 15-minute was never a constraint - it was a philosophy. And philosophies change.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The path forward is tactical:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Consolidate micro-orchestration.&lt;/strong&gt; Migrate code heavy workflows from Step Functions to Durable Functions. If your state machine is 80% ASL and 20% Lambda then flip it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Embrace AI agents as first-class workloads.&lt;/strong&gt; Durable execution is the substrate for long-running AI workflows. Checkpoint after expensive LLM calls and use &lt;code&gt;waitForCallback()&lt;/code&gt; for human-in-the-loop scenarios.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mandate determinism audits in code review.&lt;/strong&gt; Teams moving to Durable Functions need new habits. Are non-deterministic operations wrapped in steps? Is global state avoided? This intuition needs to be built.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keep Step Functions for macro-orchestration.&lt;/strong&gt; Don&apos;t throw out the visual state machine baby with the ASL bathwater. Use Step Functions for high-level infrastructure coordination and workflows that require non-technical observability. Use Durable Functions for application logic.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The future of serverless is durable. Not because AWS declared it, but because we finally have the tools to write code the way we actually think - continuously, resiliently, and without watching the clock. Let&apos;s use them....&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Lambda Durable Functions vs Step Functions: When to Use Which Workflow Tool]]></title><description><![CDATA[We've all been there - assembling IKEA furniture and can't decide whether to follow the wordless pictorial instructions or just read the…]]></description><link>https://mayankraj.com/blog/lambda-durable-vs-step-functions</link><guid isPermaLink="false">https://mayankraj.com/blog/lambda-durable-vs-step-functions</guid><pubDate>Sat, 13 Dec 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;We&apos;ve all been there - assembling IKEA furniture and can&apos;t decide whether to follow the wordless pictorial instructions or just read the text description at the bottom? The pictures are great when you need to show someone else how you built it, or even to visually trace back when you forget a step. But sometimes you just want to say &quot;attach bracket A to panel B with screws&quot;.&lt;/p&gt;
&lt;p&gt;That&apos;s the exact tension between AWS Step Functions and Lambda Durable Functions. At surface they both orchestrate workflows. They both handle long-running processes. And yet, choosing between them isn&apos;t about &quot;which one is better&quot; but instead about matching the tool to how you think and how your team operates.&lt;/p&gt;
&lt;p&gt;AWS didn&apos;t announce Durable Functions to deprecate Step Functions. They announced it because orchestration has two fundamentally different modes today: &lt;strong&gt;Macro-Orchestration&lt;/strong&gt; (connecting separate systems, requiring visual oversight) and &lt;strong&gt;Micro-Orchestration&lt;/strong&gt; (application logic that needs to live in code). Let&apos;s talk about when to use which.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-false-dilemma-will-durable-functions-replace-step-functions&quot;&gt;The False Dilemma: &quot;Will Durable Functions Replace Step Functions?&quot;&lt;/h2&gt;
&lt;p&gt;No. And asking this question misses the point entirely. Here&apos;s the mental model: Step Functions are for &lt;strong&gt;infrastructure orchestration&lt;/strong&gt;. So workflows that tie together disparate AWS services, need to be visible to non-technical stakeholders, and benefit from a graphical representation. Durable Functions instead are for &lt;strong&gt;application orchestration&lt;/strong&gt;. That&apos;s the business logic that&apos;s intrinsic to your codebase, needs to be tested locally, and doesn&apos;t want to live in separate YAML files.&lt;/p&gt;
&lt;p&gt;In my experience, the teams that struggle with this decision are the ones trying to force one tool to do the other&apos;s job. The arrival of Durable Functions doesn&apos;t obsolete Step Functions but instead it &lt;strong&gt;clarifies their roles&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-to-reach-for-step-functions-the-management-friendly-scenarios&quot;&gt;When to Reach for Step Functions: The &quot;Management-Friendly&quot; Scenarios&lt;/h2&gt;
&lt;h3 id=&quot;1-visual-observability-for-non-technical-stakeholders&quot;&gt;1. Visual Observability for Non-Technical Stakeholders&lt;/h3&gt;
&lt;p&gt;IMO this is a big plus point. We may say we love logs, but GUI beats wall of text everytime. Step Functions&apos; execution graph is a godsend at 3 AM for the on-call engineer who didn&apos;t write the code needs to see which state failed. The graphical Workflow Studio and execution timeline provide immediate insight. You literally can point at the red box and say, &quot;This is where it broke.&quot;&lt;/p&gt;
&lt;p&gt;For workflows that span multiple business units or require audit trails that executives actually look at, Step Functions&apos; visual language is absolutely worth the cost. Explaining CloudWatch Logs mean dozens of Show &amp;#x26; Tells, Runbooks for queries and what not. You don&apos;t want that. Trust me.&lt;/p&gt;
&lt;h3 id=&quot;2-low-code-aws-service-integrations&quot;&gt;2. Low-Code AWS Service Integrations&lt;/h3&gt;
&lt;p&gt;Step Functions offers direct, SDK-less integrations with well over 200 AWS services. This is where the &quot;managed service&quot; shines. Need to write an item to DynamoDB, start an ECS task, or invoke another Lambda? Step Functions can do it without any code in a Lambda function. (No &lt;code&gt;npm install&lt;/code&gt;, no dependency hell, no &quot;works on my machine&quot; debugging.)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;No cold starts. No deployment packages. Just a state machine definition pointing at AWS APIs.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is especially very powerful for infrastructure automation. Provisioning an RDS instance, waiting for it to be available, brewing coffee, running a migration script, updating Route 53, sipping that coffee - these are infrastructure tasks. They&apos;re not &quot;business logic.&quot; They&apos;re system choreography and Step Functions being in the best moves possible. They provides a clear architectural boundary: the state machine &lt;strong&gt;is&lt;/strong&gt; infrastructure, and it&apos;s managed as infrastructure.&lt;/p&gt;
&lt;h3 id=&quot;3-long-running-workflows-with-minimal-compute&quot;&gt;3. Long-Running Workflows with Minimal Compute&lt;/h3&gt;
&lt;p&gt;If your workflow is &quot;wait 7 days, send email, wait 14 days, send another email,&quot; you&apos;re not doing heavy computation between waits. Step Functions charges for state transitions, and those long waits are free. The &lt;strong&gt;$25 per million transitions&lt;/strong&gt; makes sense here because you&apos;re not transitioning much and you&apos;re mostly just waiting.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-to-reach-for-durable-functions-the-developer-first-scenarios&quot;&gt;When to Reach for Durable Functions: The &quot;Developer-First&quot; Scenarios&lt;/h2&gt;
&lt;h3 id=&quot;1-the-workflow-is-intrinsic-to-your-application-code&quot;&gt;1. The Workflow Is Intrinsic to Your Application Code&lt;/h3&gt;
&lt;p&gt;Look here if your orchestration logic involves complex data manipulation, nested loops, conditional branching based on API responses, or reliance on third-party libraries. Expressing these in native Python or Node.js is exponentially more productive than writing hundreds of lines of Amazon States Language.&lt;/p&gt;
&lt;p&gt;Let work with an example: Processing a CSV file where each row requires calling an external API and retrying on failure. Then aggregating results, and finally writing to multiple destinations based on the aggregated data. In Step Functions, that&apos;s a Choice state, a Map state, error handling with Retry and Catch, and probably a few Lambda functions anyway. In Durable Functions, it&apos;s a &lt;code&gt;for&lt;/code&gt; loop with a &lt;code&gt;try/catch&lt;/code&gt; and some &lt;code&gt;context.step()&lt;/code&gt; calls. &lt;strong&gt;Which one do you want to maintain?&lt;/strong&gt;&lt;/p&gt;
&lt;h3 id=&quot;2-local-debugging-and-unit-testing&quot;&gt;2. Local Debugging and Unit Testing&lt;/h3&gt;
&lt;p&gt;This is what makes me excited about Durable Functions. It&apos;s the &quot;holy grail&quot; of serverless development that Step Functions never delivered. With Durable Functions, we can:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Mock the &lt;code&gt;DurableContext&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Write standard Jest or Pytest tests&lt;/li&gt;
&lt;li&gt;Step through the orchestration logic in our IDE WITH BREAKPOINTS !&lt;/li&gt;
&lt;li&gt;Run the entire workflow locally without deploying to AWS - I repeat - Run and not Emulate&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;No more &quot;deploy and pray&quot; debugging where you push changes to AWS and watch the CloudWatch Logs to see if it worked. The feedback loop goes from minutes to seconds.&lt;/p&gt;
&lt;h3 id=&quot;3-high-frequency-loops-and-data-processing&quot;&gt;3. High-Frequency Loops and Data Processing&lt;/h3&gt;
&lt;p&gt;Remember the cost model: Step Functions charge &lt;strong&gt;$25 per million state transitions&lt;/strong&gt;. Durable Functions charge &lt;strong&gt;$8 per million durable operations&lt;/strong&gt; plus Lambda compute.&lt;/p&gt;
&lt;p&gt;For a workflow that processes 10,000 items in a loop:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Step Functions:&lt;/strong&gt; 3 state transitions per item (Task, Choice, Loop) = 30,000 transitions = &lt;strong&gt;$0.75&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Durable Functions:&lt;/strong&gt; 10,000 durable operations + minimal compute = &lt;strong&gt;~$0.10&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;4-your-team-lives-in-the-ide-not-the-aws-console&quot;&gt;4. Your Team Lives in the IDE, Not the AWS Console&lt;/h3&gt;
&lt;p&gt;For teams that treat infrastructure-as-code as a necessary evil rather than a core workflow, Durable Functions fundamentally end up eliminating the context-switching tax. The orchestration logic and the application logic are in the same file, versioned together, reviewed together ...and break together. You don&apos;t have separate Terraform files for the state machine definition. You don&apos;t have to mentally map ASL states to Lambda function names.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-hybrid-pattern-using-both-yes-really&quot;&gt;The Hybrid Pattern: Using Both (Yes, Really)&lt;/h2&gt;
&lt;p&gt;Here&apos;s the thing nobody tells you: &lt;strong&gt;you can use both in the same architecture&lt;/strong&gt;. In fact, you probably should and if you are not - you need to have strong reasons for it.&lt;/p&gt;
&lt;p&gt;Use Step Functions for the &lt;strong&gt;macro-orchestration&lt;/strong&gt;: &quot;Provision infrastructure, deploy application, run smoke tests, update DNS.&quot; These are infrastructure tasks that benefit from the visual execution graph, native AWS integration. Less things can go wrong when you don&apos;t write another line of logic.&lt;/p&gt;
&lt;p&gt;Use Durable Functions for the &lt;strong&gt;micro-orchestration&lt;/strong&gt; within those steps: &quot;Process customer data, chain multiple AI model calls, handle human-in-the-loop approval.&quot; These are application tasks that benefit from code flexibility. They are specific to your business logic, and you know the best way to handle them.&lt;/p&gt;
&lt;p&gt;A Step Functions state machine can invoke a Durable Function. The Durable Function does its complex logic, checkpoints its progress, and returns a result. Step Functions continues with the next infrastructure task. You get the best of both worlds: visual infrastructure orchestration and flexible application logic.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-takeaway-two-tools-one-architecture&quot;&gt;The Takeaway: Two Tools, One Architecture&lt;/h2&gt;
&lt;p&gt;AWS didn&apos;t cave into the Durable Functions to create confusion. It well timed, and I think well thought of. We got it because orchestration has evolved into two very distinct modes today and we&apos;ve been forcing one tool to do both jobs for a decade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Step Functions&lt;/strong&gt; excel at infrastructure orchestration that needs visual debugging and cross-team visibility. Keep using them for that. They&apos;re purpose-built for it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Durable Functions&lt;/strong&gt; excel at application logic that needs code flexibility and local testability. Start using them for that. They&apos;re purpose-built for it.&lt;/p&gt;
&lt;p&gt;The maturity of serverless isn&apos;t about replacing old tools with new ones. It&apos;s about having the right tool for each job and also knowing when to reach for which. Let&apos;s build architectures that match how we think, not how we&apos;re forced to think by our tools.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Library vs. The Warehouse: Why Your Database Can't Have It All]]></title><description><![CDATA[Your application just hit 10,000 writes per second, and your database is quite literally screaming. Looking at the write latency chart is…]]></description><link>https://mayankraj.com/blog/btree-lsm-database-storage-tradeoffs</link><guid isPermaLink="false">https://mayankraj.com/blog/btree-lsm-database-storage-tradeoffs</guid><pubDate>Mon, 17 Nov 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Your application just hit 10,000 writes per second, and your database is quite literally screaming. Looking at the write latency chart is like watching a heart attack ...while it&apos;s in progress! Your on-call engineer is panic texting, ALL CAPS, in the team chat. Someone just spoke out the obvious - &quot;maybe we need to throw more RAM at it.&quot; You&apos;ve been here before, many many times. Every engineering team eventually faces the same brutal question: do we optimize for reading data, or writing it?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;We can&apos;t have both.&lt;/strong&gt; Not really.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-physics-problem-or-why-libraries-and-warehouses-dont-work-the-same-way&quot;&gt;The Physics Problem (Or: Why Libraries and Warehouses Don&apos;t Work the Same Way)&lt;/h2&gt;
&lt;p&gt;Let&apos;s talk about two very different buildings. The first is a traditional library, the kind with the musty smell and the Dewey Decimal System. Every book has exactly one location, known neighbors and predictable rack. When a new book comes in, a librarian walks to a very precise shelf, shifts everything over (yes, physically moves books), and slots it into its ordained position. Want to get that book later? Instant. You look it up in the catalog, walk directly to QA76.76, and voila - there it is. The second building is an Amazon fulfillment center. Packages arrive by the truckload. Workers don&apos;t, or rather can&apos;t carefully organize them by the category which means that they throw them into the nearest available bin as fast as humanly possible. Need to find something? You end up having to scan barcodes, check multiple bins, and hope your &quot;smart system&quot; knows which bins might contain your item. But the receive rate? Absolutely insane.&lt;/p&gt;
&lt;p&gt;B-Trees are the library while LSM Trees are the warehouse.&lt;/p&gt;
&lt;p&gt;Every modern database is exactly one of these two buildings. Even worse, few act as the clever hybrid trying to be both, usually succeeding at neither. The fundamental constraint isn&apos;t in the software but the physics. Specifically, the physics of moving bits between RAM and disk.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-b-tree-pay-now-read-fast-forever&quot;&gt;The B-Tree: Pay Now, Read Fast Forever&lt;/h2&gt;
&lt;p&gt;B-Tree have dominated relational databases since the 1970s and for a good reason: they minimizes the number of times you have to ask the disk for data. Back then hard drives were physical entities with actual spinning platters and moving read/write heads. A random disk seek could take 10 milliseconds. That&apos;s couple of centuries in human timescale. In the time it takes to move a mechanical arm across a platter, a modern CPU could execute about 30 million instructions. We&apos;re looking at ~100x performance gap between sequential and random I/O on HDDs.&lt;/p&gt;
&lt;p&gt;The B-Tree was designed for this mechanical sloth as per today&apos;s standards. It organizes data into fixed-size pages (~4-16 KB) and maintains a sorted tree structure where the internal nodes act as the navigation layer. The tree&apos;s fan-out is high wherein each internal node can point to hundreds of children. This enables databases with millions of records to remains only 3 or 4 levels deep. What you get in return is a point lookup requiring just a few disk reads, most of which are probably cached in memory anyway. Range scans are even better. Given that the leaf nodes are linked together, once you find your starting point, you just walk forward reading contiguous pages. The database is &lt;em&gt;always&lt;/em&gt; organized, &lt;em&gt;always&lt;/em&gt; sorted, &lt;em&gt;always&lt;/em&gt; ready for your read query.&lt;/p&gt;
&lt;h3 id=&quot;the-write-tax&quot;&gt;The Write Tax&lt;/h3&gt;
&lt;p&gt;Writes are where the librarian starts sweating. When you update a single 100 byte record in the B-Tree, you don&apos;t just write 100 bytes to disk. You instead write the whole of 16 KB page containing that record. That&apos;s 160x more data than you actually changed. It gets even worse. To ensure durability, every modification is written twice: once to a Write Ahead Log (WAL), and once to the B-Tree page itself. So now we&apos;re at 320x amplification for that poor 100-byte update.&lt;/p&gt;
&lt;p&gt;And then there&apos;s fragmentation as B-Tree nodes are rarely 100% full. If they were then every insertion would trigger an expensive page split! In practice, a B-Tree under an active workload might have a 50-70% fill factor. That means 30-50% of your disk space is just... well... empty. Reserved for future insertions which may never come.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The B-Tree is the database equivalent of paying your taxes upfront. You may suffer during the write, but your reads never have to worry.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-lsm-tree-the-art-of-organized-procrastination&quot;&gt;The LSM Tree: The Art of Organized Procrastination&lt;/h2&gt;
&lt;p&gt;In 1996 &lt;a href=&quot;https://en.wikipedia.org/wiki/Patrick_O%27Neil&quot;&gt;Patrick O&apos;Neil&lt;/a&gt; (aka DB legend) and colleagues asked a dangerous question: &lt;em&gt;What if we just... didn&apos;t organize data immediately?&lt;/em&gt; The Log-Structured Merge Tree, or LSM Tree was born from this very simple question. Sequential writes are 100x faster than random writes on HDDs. If we stop trying to update data in place and instead treat every write as an append operation, we can match the theoretical maximum throughput of the storage device, at least for the writing part.&lt;/p&gt;
&lt;p&gt;Here&apos;s how it works. When your application writes a record:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Append to the Write-Ahead Log&lt;/strong&gt;: The record is appended to a sequential log file on disk. This is extremely fast because we&apos;re just writing to the end of a file. There&apos;s no seeking, no shuffling, just raw sequential bandwidth.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Insert into the MemTable&lt;/strong&gt;: The record is added to an in-memory sorted structure. This happens in RAM in nanoseconds.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Flush to Disk&lt;/strong&gt;: When the MemTable fills up (say 64 MB), it&apos;s marked immutable and flushed to disk as a new Sorted String Table (SSTable). This flush is purely sequential and so at maximum speed.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Immutable writes&lt;/strong&gt;: SSTable is strictly immutable. So If you update a record then you don&apos;t modify the old one but you write a new version. The old version sits there until a background process cleans it up later.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The throughput difference is staggering. LSM Trees like RocksDB can sustain 2-5x higher write rates than B-Trees. In Facebook&apos;s benchmarks, RocksDB regularly exceeded 100 MB/s ingestion while WiredTiger (a B-Tree engine) bottlenecked at 20 MB/s.&lt;/p&gt;
&lt;h3 id=&quot;the-read-debt&quot;&gt;The Read Debt&lt;/h3&gt;
&lt;p&gt;You cannot escape tradeoffs. The warehouse model has a problem and it&apos;s in finding the stuff you stored. To read a single key in an LSM Tree, you have to check:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The active MemTable (in RAM)&lt;/li&gt;
&lt;li&gt;Any immutable MemTables waiting to be flushed&lt;/li&gt;
&lt;li&gt;All the Level 0 SSTables on disk with overlapping key ranges&lt;/li&gt;
&lt;li&gt;One SSTable per level in the hierarchy&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As you can imagine, this is slow. A single key lookup could require checking a dozen files across RAM and disk. So LSM Trees cheat in two ways:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bloom Filters&lt;/strong&gt;: Before reading an SSTable the DB engine checks a probabilistic data structure to see if the key is definitely not in this file? If the Bloom filter says no, then it skips the entire file or else it&apos;ll have to check. There are chances of false positive, but majority cases proceeds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sparse Indexes&lt;/strong&gt;: Inside each SSTable, data is stored in sorted blocks. With this the system maintains a sparse index with pointers indicating the start and end keys for each block. This allows binary search within the file, narrowing down to a single 4 KB block instead of scanning the entire thing.&lt;/p&gt;
&lt;p&gt;Even with these tricks, LSM&apos;s read latency is higher than B-Trees. The 95th percentile latency in production LSM systems can be 2-3x worse than B-Trees for point lookups.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-nvme-twist&quot;&gt;The NVMe Twist&lt;/h2&gt;
&lt;p&gt;For decades the LSM Tree&apos;s dominance in write-heavy workloads was justified by the extreme cost of random writes on spinning disks. But NVMe SSDs have fundamentally changed the equations now. Modern NVMe drives don&apos;t have any mechanical heads. They use parallel flash channels which can handle thousands of random 4 KB writes per second at microsecond latency. The performance gap between random and sequential I/O has collapsed to as little as 2-3x. Now as the random writes are no longer prohibitively expensive, the value add to the LSM Tree&apos;s weakens.&lt;/p&gt;
&lt;p&gt;Recent research on Bf-Trees shows write throughput 6x higher than traditional B-Trees and point lookups 2x faster than both B-Trees and LSM Trees! It&apos;s a next-generation B-Tree variant designed for SSDs. The future of database storage might not be LSM Trees and we might get back to the roots of B-Trees.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-framework-organize-now-or-organize-later&quot;&gt;The Framework: Organize Now, or Organize Later?&lt;/h2&gt;
&lt;p&gt;You&apos;ll learn within years of working with databases - there is no perfect database. The space is as specialized as it gets. Each has its tradeoffs and is applicable for a given use case.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;B-Trees organize at ingestion.&lt;/strong&gt; Write pays the cost of finding the correct sorted position then updating pages in place, and finally maintaining the tree invariant. The benefit? Reads are orders of magnitude faster.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;LSM Trees organize at compaction.&lt;/strong&gt; Every write feels fast because you&apos;re just appending. The cost is deferred to background processes that merge, sort, and garbage-collect. The benefit? Extreme write throughput.&lt;/p&gt;
&lt;p&gt;This isn&apos;t just a software problem but it&apos;s hardware and physics. The RUM conjecture (Read, Update, Memory overhead) states that you can optimize for two of these three dimensions, but not all three. B-Trees choose Read and Memory efficiency at the cost of Update performance. LSM Trees choose Update efficiency at the cost of Read complexity and Space overhead. The library and the warehouse both work. It&apos;s just that they are optimized for different realities. Understand the trade-offs well enough that when your database inevitably melts down at 3 AM, you&apos;ll know why and more importantly what lever to pull.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[DynamoDB Single Table Design: A Brilliant Hack That's Probably Torturing Your Team]]></title><description><![CDATA[I'll be the brave one and say it out loud - Single Table Design in DynamoDB is premature optimization for 90% of applications. Now before…]]></description><link>https://mayankraj.com/blog/dynamodb-single-table-design-premature-optimization</link><guid isPermaLink="false">https://mayankraj.com/blog/dynamodb-single-table-design-premature-optimization</guid><pubDate>Tue, 21 Oct 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;I&apos;ll be the brave one and say it out loud - Single Table Design in DynamoDB is premature optimization for 90% of applications.&lt;/p&gt;
&lt;p&gt;Now before you close this tab in righteous indignation, just hear me out. I&apos;ve seen this pattern implemented beautifully at hyper scale ...and I&apos;ve also seen it turn a $200/month database into a $50,000/year developer productivity black hole. The difference isn&apos;t about the pattern itself. &lt;strong&gt;It&apos;s about whether your application actually needs it.&lt;/strong&gt; And in most of the cases - it doesn&apos;t.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-filing-cabinet-thought-experiment&quot;&gt;The Filing Cabinet Thought Experiment&lt;/h2&gt;
&lt;p&gt;Imagine you&apos;ve just been hired at a company. Your first task: find Invoice #4521 from March.&lt;/p&gt;
&lt;p&gt;In &lt;strong&gt;Office A&lt;/strong&gt;, there are three filing cabinets: &quot;Invoices,&quot; &quot;Customer Records,&quot; and &quot;Orders.&quot; You walk to the Invoices cabinet, flip to March, pull the file. Done. Your predecessor labeled everything clearly. Why? Because they weren&apos;t a psychopath.&lt;/p&gt;
&lt;p&gt;In &lt;strong&gt;Office B&lt;/strong&gt;, there&apos;s one gigantic filing cabinet. Everything is in there - labelled with postits - but all of it in there. Invoices, customer records, orders, important notes from 2019, someone&apos;s lunch receipt. Each file is labeled with a code: &lt;code&gt;INVOICE#4521&lt;/code&gt;, &lt;code&gt;CUSTOMER#789&lt;/code&gt;, &lt;code&gt;ORDER#2024-03-15#001&lt;/code&gt;. The robot filing clerk who designed this system can retrieve any document in 3 milliseconds. But not you - you don&apos;t know the invoice number for the March recept, and so you have to look through each-and-every-file-one-by-one.&lt;/p&gt;
&lt;p&gt;This is Single Table Design versus Multi-Table Design. And that 3-millisecond advantage? In my experience, it&apos;s rarely worth the tears.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;a-brief-history-of-a-brilliant-hack&quot;&gt;A Brief History of a Brilliant Hack&lt;/h2&gt;
&lt;p&gt;This was a shocker for me - single Table Design wasn&apos;t born from best practices but was born from desperation.&lt;/p&gt;
&lt;p&gt;In the early 2010s, DynamoDB was a far more constrained beast. It was and always been a specialised tool, for a specific job. Back then you got five Global Secondary Indexes (GSIs) per table. A hard limit of five. If you had six different query patterns, too bad. Spin up another table, provision separate capacity units, and watch your AWS bill as well as the operational overhead multiply.&lt;/p&gt;
&lt;p&gt;A small group of architects at Amazon (led by Rick Houlihan) found a clever workaround. Instead of fighting the limitations, they embraced them. They created &quot;Index Overloading&quot;: generic partition keys (&lt;code&gt;PK&lt;/code&gt;) and sort keys (&lt;code&gt;SK&lt;/code&gt;) that could store any entity type. A single GSI could now support users, orders, invoices, and product catalogs. The robot filing clerk was happy.&lt;/p&gt;
&lt;p&gt;A bug was officially now a feature. This was genuinely brilliant. In 2012.&lt;/p&gt;
&lt;p&gt;Here&apos;s what changed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GSI limits increased from 5 to 20 to 25&lt;/li&gt;
&lt;li&gt;On-Demand pricing eliminated the &quot;table-per-entity&quot; cost penalty&lt;/li&gt;
&lt;li&gt;DynamoDB got faster, and networks got faster too&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The original constraints that justified Single Table Design? &lt;strong&gt;Largely gone.&lt;/strong&gt; There&apos;s no reason for the pattern to exist but it still does. It does not because it is still optimal, but because it has become dogma. Conference talks, blog posts, and certification exams continued to preach it as &quot;the way.&quot; I&apos;m literary adding to the debt here with this post!&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-trade-off-nobody-talks-about&quot;&gt;The Trade-Off Nobody Talks About&lt;/h2&gt;
&lt;p&gt;Let&apos;s get mechanical for a moment. Single Table Design offers one undeniable advantage: &lt;strong&gt;pre-joined data retrieval.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;By cramming a user details, their recent orders, their shipping addresses and what not into the same partition, you can fetch everything in a single query. One network round-trip instead of potential dozens. In theory, this is fantastic. The database is doing the work.&lt;/p&gt;
&lt;p&gt;In practice, you&apos;ve just shifted that work (and then some) to your application code.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Single Table&lt;/th&gt;
&lt;th&gt;Multi-Table&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Query Speed&lt;/td&gt;
&lt;td&gt;Marginally faster (single round-trip)&lt;/td&gt;
&lt;td&gt;Slightly slower (parallel queries)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code Complexity&lt;/td&gt;
&lt;td&gt;High (maintain composite keys, prefixes)&lt;/td&gt;
&lt;td&gt;Low (standard CRUD operations)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;Nightmare (generic &lt;code&gt;PK&lt;/code&gt;, &lt;code&gt;SK&lt;/code&gt; everywhere)&lt;/td&gt;
&lt;td&gt;Straightforward (descriptive names)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema Changes&lt;/td&gt;
&lt;td&gt;Super Expensive ETL migrations&lt;/td&gt;
&lt;td&gt;Add a table, add an index&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Analytics Readiness&lt;/td&gt;
&lt;td&gt;Painful (everything&apos;s mixed)&lt;/td&gt;
&lt;td&gt;Natural (one table = one concept)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The database is slightly more happier. Your developers - significantly less happy. And here&apos;s the thing about developer happiness: it directly correlates with how fast your team ships features and how often they introduce bugs.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-cognitive-load-crisis&quot;&gt;The Cognitive Load Crisis&lt;/h2&gt;
&lt;p&gt;Forrest Brazeal once described a well-optimized Single Table Design as looking &quot;more like machine code than a simple spreadsheet.&quot; He wasn&apos;t being hyperbolic.&lt;/p&gt;
&lt;p&gt;Here&apos;s what debugging looks like in a Single Table:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;PK                    SK                          Data
USER#123              METADATA#                   { &quot;name&quot;: &quot;Alice&quot;, &quot;email&quot;: &quot;...&quot; }
USER#123              ORDER#2024-03-15#001        { &quot;total&quot;: 59.99, ... }
USER#123              ORDER#2024-03-14#002        { &quot;total&quot;: 124.50, ... }
USER#123              ADDRESS#HOME                { &quot;street&quot;: &quot;123 Main St&quot;, ... }
USER#123              ACTIVITYLOG#2024-03-15T10:  { &quot;action&quot;: &quot;login&quot;, ... }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Quick: find all the orders ...and filter by date. Now explain to the new hire why the sort key starts with &lt;code&gt;ORDER#&lt;/code&gt; but the address one starts with &lt;code&gt;ADDRESS#&lt;/code&gt;. Now explain the tilde hack (appending &lt;code&gt;~&lt;/code&gt; to sort keys so parent items sort correctly)...&lt;/p&gt;
&lt;p&gt;This is what I call &lt;strong&gt;ASCII sorcery&lt;/strong&gt;: clever tricks that work beautifully for the database but create tribal knowledge dependencies. This is rarely documented, or even if it is - the documentation rarely keeps up with how the table grows. New team members can&apos;t be productive until they&apos;ve absorbed months of context. Simple debugging sessions become archaeology expeditions. You cannot do a &lt;code&gt;Distinct&lt;/code&gt; query to get all variation. Above all - every exporation is very very costly. Single tables are not self documented.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&quot;Don&apos;t torture your developers for $5 a month.&quot;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That&apos;s not my quote. It&apos;s the emerging consensus from engineers who&apos;ve lived through the Single Table experience.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-5-vs-5000-math&quot;&gt;The $5 vs. $5,000 Math&lt;/h2&gt;
&lt;p&gt;Let me make this concrete with some napkin math.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Single Table pitch:&lt;/strong&gt; By pooling all entities into one table, you share provisioned capacity. Economics of scale - less waste, better utilization, lower bill.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The reality:&lt;/strong&gt; On-Demand pricing charges per request, not per table. Whether your reads hit one table or twenty, the per-RRU cost is identical.&lt;/p&gt;
&lt;p&gt;But let&apos;s say you&apos;re on Provisioned Capacity, and Single Table genuinely saves you $100/month in capacity pooling efficiency. Sounds great, right?&lt;/p&gt;
&lt;p&gt;Now consider:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;40 hours of architect time designing the Single Table schema&lt;/li&gt;
&lt;li&gt;20 hours of documentation (because nobody can understand it otherwise)&lt;/li&gt;
&lt;li&gt;8 hours per bug debugging through cryptic partition keys&lt;/li&gt;
&lt;li&gt;16 hours for every schema change requiring an ETL migration&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At even $75/hour (conservative for senior engineers), you&apos;ve spent $6,000 before you&apos;ve saved your first dollar. Break-even is five years away. And that assumes no bugs, no churn, and no feature changes. (Spoiler: these assumptions have never been true for any project, ever.) You may have few blog posts, talks to speak of the design but the original goal remains a gaol for the future.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-analytics-wall&quot;&gt;The Analytics Wall&lt;/h2&gt;
&lt;p&gt;Every application eventually needs to answer adhoc questions like:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&quot;What was our year-over-year growth by product category?&quot;&lt;/li&gt;
&lt;li&gt;&quot;Which users are most likely to churn?&quot;&lt;/li&gt;
&lt;li&gt;&quot;What&apos;s the average order value by region?&quot;&lt;/li&gt;
&lt;li&gt;&quot;Which user caused the spike in April?&quot;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;DynamoDB is extremely exceptional at transactional workloads ...but also equally poor for analytics. Asking DynamoDB to do analytics is like asking a sprinter to run a marathon: technically possible, deeply uncomfortable for everyone involved. When you inevitably need to stream your data into Redshift or a BI tool, Single Table Design becomes your worst enemy.&lt;/p&gt;
&lt;p&gt;Remember that beautiful mixed table of users, orders, and addresses? Your analytics team now sees:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;pk&lt;/th&gt;
&lt;th&gt;sk&lt;/th&gt;
&lt;th&gt;value (SUPER type)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;USER#123&lt;/td&gt;
&lt;td&gt;METADATA#&lt;/td&gt;
&lt;td&gt;{&quot;name&quot;: &quot;Alice&quot;, ...}&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;USER#123&lt;/td&gt;
&lt;td&gt;ORDER#2024-03-15#001&lt;/td&gt;
&lt;td&gt;{&quot;total&quot;: 59.99, ...}&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;They&apos;ll need to write PartiQL queries to &quot;shred&quot; this semi-structured data back into relational format. The performance gains you realized in your transactional database? You have to give it back, with interest, in the analytical layer.&lt;/p&gt;
&lt;p&gt;Multi-Table Design, by contrast, maps naturally: &lt;code&gt;Users&lt;/code&gt; table → &lt;code&gt;users&lt;/code&gt; in Redshift. &lt;code&gt;Orders&lt;/code&gt; table → &lt;code&gt;orders&lt;/code&gt; in Redshift. Your analysts can write normal SQL on day one.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-microservices-trap&quot;&gt;The Microservices Trap&lt;/h2&gt;
&lt;p&gt;Single Table Design and microservices are philosophical opposites pretending to be friends.&lt;/p&gt;
&lt;p&gt;Microservices is built on each service having its own data. Independent deployment, independent scaling, clear separation boundaries. Single Table instead forces multiple services to share one table. A schema change for Service A can break Service B. A &quot;noisy neighbor&quot; service can end up throttling everyone. IAM permissions become a nightmare (good luck granting access to specific entity types in a shared table).&lt;/p&gt;
&lt;p&gt;This is the &lt;strong&gt;Shared Database Anti-Pattern&lt;/strong&gt; wearing a NoSQL costume. The whole point of separate services is reducing blast radius. A single shared table instead makes it a ticking clock.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-single-table-actually-makes-sense-the-10&quot;&gt;When Single Table Actually Makes Sense (The 10%)&lt;/h2&gt;
&lt;p&gt;I&apos;m not arguing Single Table Design is always wrong. In specific contexts, it&apos;s the right tool:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Hyper-scale core services:&lt;/strong&gt; If you&apos;re processing millions of queries per second and operate at sub-millisecond latency which directly impacts revenue, the eliminated network round-trip matters. Think identity services, high-frequency trading systems, or core infrastructure at Netflix-scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Stable, immutable access patterns:&lt;/strong&gt; URL shorteners, logging pipelines, or event stores where access patterns that are genuinely set in stone. If your queries haven&apos;t changed in three years and won&apos;t change in the next three, Single Table&apos;s rigidity is a feature to embrace.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Global Database:&lt;/strong&gt; DynamoDB was one of the first truely globally available database I worked with. Regional database, across the globe, all synced with each other - it was magical. Understanding the consistency model was the key to making it work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Expert teams with DynamoDB depth:&lt;/strong&gt; If your team has shipped multiple Single Table implementations and genuinely understands the trade-offs, you&apos;re equipped to make an informed choice.&lt;/p&gt;
&lt;p&gt;For everyone else (startups iterating on product-market fit, teams with normal DynamoDB experience, applications that will need analytics, systems where requirements change quarterly), Multi-Table Design is almost certainly the better bet.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-pragmatic-path-forward&quot;&gt;The Pragmatic Path Forward&lt;/h2&gt;
&lt;p&gt;Whenever I hear DynamoDB, or see a Single Table design doc - the following goes through my mind:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can we start with Multi-Table Design?&lt;/strong&gt; One table per entity type, with descriptive attribute names. Standard GSIs for your query patterns. This gives us readable data in the AWS Console and straightforward debugging. At early stage, things would change, access patterns would be refined and data driven decissions need analytics layer with flexible querying.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can we use parallel queries?&lt;/strong&gt; DynamoDB is fast. But then two parallel queries at 15ms each often complete faster than developers expect — and the scaffolding is dramatically simpler. When you don&apos;t need to shave off milliseconds - don&apos;t try to.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;I&apos;d consider Single Table only when:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We&apos;ve profiled the application and network latency is definately the bottleneck&lt;/li&gt;
&lt;li&gt;The access patterns are stable and well-understood&lt;/li&gt;
&lt;li&gt;The team has the expertise to maintain it, or we&apos;ve experimented with non-critical use case first&lt;/li&gt;
&lt;li&gt;The math actually works (savings &gt; complexity cost)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Make use of tooling if you must.&lt;/strong&gt; There are now libraries like ElectroDB, DynamoDB-Toolbox, and OneTable which can reduce Single Table complexity. They&apos;re good tools, genuinely. But they&apos;re also treating a symptom rather than questioning whether you need the disease.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-bottom-line&quot;&gt;The Bottom Line&lt;/h2&gt;
&lt;p&gt;Single Table Design was a brilliant architectural hack for a database service which had real limitations. Today those limitations have largely been removed. The pattern remains useful in specific large scale scenarios, but for the vast majority of applications - it sign of premature optimization of the most expensive kind: trading marginal database efficiency for massive developer inefficiency.&lt;/p&gt;
&lt;p&gt;The most valuable resource in your organization isn&apos;t the RCU or the WCU but instead it&apos;s the focused, productive time of your engineering team. Go build for agility first. Optimize for database efficiency only when the profiler tells you to and not when a conference talk from 2018 implies you should.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Cell-Based Architecture: When Isolation Beats Integration]]></title><description><![CDATA[On June 30, 2021, Slack went dark. Not completely - that would have been easier to explain. Instead the platform entered a state we…]]></description><link>https://mayankraj.com/blog/cell-based-architecture-blast-radius-containment</link><guid isPermaLink="false">https://mayankraj.com/blog/cell-based-architecture-blast-radius-containment</guid><pubDate>Wed, 10 Sep 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;On June 30, 2021, Slack went dark. Not completely - that would have been easier to explain. Instead the platform entered a state we engineers dread more than a clean outage: the partial failure. Turns out that a single AWS Availability Zone was dropping packets. Just a few percentage points of loss. But in a distributed system where one user action fans out into hundreds of internal API calls, effect of these dropped packets propagated like poison through the bloodstream. Within minutes, users across all three availability zones couldn&apos;t send messages. The failure of one data center had somehow killed the entire platform.&lt;/p&gt;
&lt;p&gt;The modern distributed systems are not like traditional applications wherein they crash cleanly when something breaks. They&apos;re more like submarines wherein the hull is divided into watertight compartments called bulkheads. So if one section floods, you seal it off and the rest of the vessel to survive. Without these compartments, a single breach sinks the ship. (famously Titanic had compartments, but not enough of them, and not sealed at the top. One breach became five, then fifteen...)&lt;/p&gt;
&lt;p&gt;Traditional microservices architectures lack these bulkheads. I&apos;ll give you that you&apos;ve split your monolith into dozens of services, each with its own codebase, it&apos;s own deployment pipeline etc. But scratch the surface and you&apos;ll find they all share the same database cluster, or maybe the same Redis cache, the same message queue, are in the same AZ ...you get the point. They&apos;re technically separate services living in a shared fate environment. When one service misbehaves, it sends a poison pill request, triggers a cascading retry storm, or simply experiences elevated latency. The blast radius is 100%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cell-based architecture is the antidote to shared fate.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-blast-radius-problem&quot;&gt;The Blast Radius Problem&lt;/h2&gt;
&lt;p&gt;Let&apos;s get on the same page first. Blast radius is the scope of impact when something fails. In traditional thinking, we measured this spatially: how many servers went down, how many users were affected etc. Modern availability engineering adds a temporal dimension: how long did the damage last. Reducing an incident from one hour to one minute is the new and imo, the more accurate blast radius reduction, even if the same number of users were momentarily impacted.&lt;/p&gt;
&lt;p&gt;Most distributed systems are designed to prevent failure. That&apos;s an unattainable goal. At any decent scale when you&apos;re serving millions of requests per second across dozens of microservices, failure is mathematically inevitable. Software bugs will slip through code review - no matter how many extra LLMS babysit PRs. Infrastructure will eventually experience &quot;gray failures&quot; where hardware degrades but doesn&apos;t cleanly die. A single tenant will send a request pattern that crashes your carefully load tested service.&lt;/p&gt;
&lt;p&gt;The correct goal is &lt;strong&gt;containment&lt;/strong&gt;. Not preventing the breach, but limiting how far the water spreads.&lt;/p&gt;
&lt;p&gt;Cell-based architecture achieves this by partitioning your entire system and not just the code, but also the compute, storage, and all dependencies. This is broken into independent, self-sufficient units called &quot;cells&quot; (or &quot;pods&quot; or &quot;shards&quot; or &quot;my-mini-app-1&quot;). Each cell is a complete miniature replica of your entire stack. If you have 100 cells and one fails, only 1% of your users are affected. The failure is isolated and contained. It&apos;s now &lt;strong&gt;Mathematically predictable.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If your blast radius is &lt;code&gt;1/N&lt;/code&gt; where &lt;code&gt;N&lt;/code&gt; is your cell count, reliability becomes a scaling problem, not an engineering breakthrough event.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-shared-infrastructure-becomes-shared-fate&quot;&gt;When Shared Infrastructure Becomes Shared Fate&lt;/h2&gt;
&lt;p&gt;Here&apos;s what Slack learned the hard way. In a horizontally scaled microservices architecture, your service instances are broken down and distributed across multiple availability zones primarily for redundancy. The frontend services in Zone A can still talk to backend services in Zone B or Zone C. Databases are replicated and synced across zones. From a resource utilization perspective, this is optimal - every server can serve any request. You&apos;ve maximized your capacity. Redundancy is on point as well.&lt;/p&gt;
&lt;p&gt;But you&apos;ve also created a perfectly connected graph. The failure spreads like wildfire When even one node in that graph starts zoning out just enough to cause timeouts but not enough to trigger circuit breakers cleanly. Your frontend in Zone A tries to call a backend in the degraded Zone B, times out, retries, and now that frontend is slow. Your load balancer keeps routing new requests to the slow frontend because it&apos;s technically still responding (just late). Users see errors. Your monitoring dashboards start lighting up like a Christmas tree. And the root cause? It&apos;s happening in a completely different availability zone than where the symptoms appear.&lt;/p&gt;
&lt;p&gt;Contrast this with the cellular model. In Slack&apos;s post incident redesign, we implemented &quot;AZ siloing.&quot; Traffic enters a specific availability zone and just stays there. Services in a Zone only talk to other services within the same Zone. The backend databases are all partitioned with each zone has its own dedicated slice of data. If Zone B starts dropping packets, the router simply stops sending traffic there. Users assigned to cells in Zone A and Zone C continue working normally. The blast radius is contained to one zone. In case case a zone failure can only ever affect 33% of the users.&lt;/p&gt;
&lt;p&gt;The system is &lt;em&gt;designed to lose a third of itself&lt;/em&gt; and keep running. That&apos;s not resilience through redundancy but it&apos;s resilience through isolation.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-anatomy-of-a-cell&quot;&gt;The Anatomy of a Cell&lt;/h2&gt;
&lt;p&gt;A cell is more than just a logical grouping as it acts like a &lt;strong&gt;hard isolation boundary&lt;/strong&gt;. Let&apos;s break down the components to get a better picture fo it:&lt;/p&gt;
&lt;h3 id=&quot;1-the-cell-router-the-thin-front-door&quot;&gt;1. The Cell Router (The Thin Front Door)&lt;/h3&gt;
&lt;p&gt;Every request enters through a cell router which is a lightweight, stateless service whose only job is to answer one question: &quot;Which cell should handle this request?&quot; The router uses a partition key (tenant ID, user ID, resource ID, geographic region etc) to deterministically map the request to a specific cell. This mapping is cached locally at the router, so doesn&apos;t add to the latency in any significant manner.&lt;/p&gt;
&lt;p&gt;The router must be dumb. If you start adding stuff like authentication logic, rate limiting, or complex business rules into the router, you&apos;ve just created a new single point of failure. The router fails, every cell is unreachable. Game over.&lt;/p&gt;
&lt;h3 id=&quot;2-the-cell-itself-the-bulkhead&quot;&gt;2. The Cell Itself (The Bulkhead)&lt;/h3&gt;
&lt;p&gt;Inside the cell is a complete self contained deployment of your application. Remember that Duplication is the key here. All the microservices, all the databases, all the caches ...even the manually added &quot;test IAM&quot; roles. This rightfully feels wasteful at first especially from a pure cost perspective. But a database connection pool leak in Cell 7 cannot affect Cell 23. A cache stampede in Cell 42 doesn&apos;t cascade to Cell 38.&lt;/p&gt;
&lt;p&gt;Cells should be designed to be &lt;strong&gt;unaware of each other&lt;/strong&gt;. This means no cross cell API calls. No shared databases. No &quot;just this one global Redis instance for session state.&quot; If two tenants in different cells need to communicate, that happens asynchronously via event buses or message queues but never ever synchronously on the hot path of a user request.&lt;/p&gt;
&lt;h3 id=&quot;3-the-control-plane-the-orchestrator&quot;&gt;3. The Control Plane (The Orchestrator)&lt;/h3&gt;
&lt;p&gt;Cells don&apos;t deploy themselves. You need a control plane that manages the lifecycle of cells. It&apos;s very very important that every cell is identical to each other. This control plane includes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Provisioning&lt;/strong&gt;: Spinning up new cells as you grow&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Health monitoring&lt;/strong&gt;: Detecting when a cell is degraded&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deployment orchestration&lt;/strong&gt;: Rolling out updates in waves (canary, blue-green etc)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Graceful evacuation&lt;/strong&gt;: Draining traffic from a sick cell without dropping the in-flight requests&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At Salesforce we have the Hyperforce platform, which uses a control plane that manages deployments across hundreds of cells globally for thousands of services. It&apos;s equipped with automated rollback, deployment strategies and what not. The control plane itself must be highly available, but it operates out-of-band from the data plane - it&apos;s not in the critical path of user requests.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;shuffle-sharding-isolation-on-steroids&quot;&gt;Shuffle Sharding: Isolation on Steroids&lt;/h2&gt;
&lt;p&gt;Now that we have isolated cells - we can make better use of them. Standard cell based architecture gives you &lt;code&gt;1/N&lt;/code&gt; blast radius i.e. If you have 100 cells and one fails, 1% of your users are affected. But if you&apos;re operating a multi-tenant SaaS platform every tenant on the cell has equal power to poison pill every other tenant on that cell. Say they send a query that triggers a bug and crashes the services in their cell. With standard sharding, every tenant in that cell goes down together.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shuffle sharding&lt;/strong&gt; aims to solves this with combinatorics. Instead of assigning each tenant to a single cell, you assign them to a virtual shard composed of multiple randomly selected workers from the global fleet. Think of it like dealing a hand of cards from a shuffled deck—each tenant gets their own unique hand. The math is beautiful for this. In a cluster of 100 workers, if you assign each tenant a virtual shard of 5 workers, the number of unique combinations is:&lt;/p&gt;
&lt;p&gt;$$
\binom{100}{5} = \frac{100!}{5!(100-5)!} = 75,287,520
$$&lt;/p&gt;
&lt;p&gt;If one tenant crashes their 5 assigned workers, the probability that another tenant shares those exact same 5 workers is ...well ...effectively zero. Even if two tenants share one or two workers (which is statistically likely), each tenant retains 60-80% of their capacity albeit degraded, but still operational.&lt;/p&gt;
&lt;p&gt;You want proof that this works? AWS uses this technique for Route 53. When you create a hosted zone, you&apos;re not assigned to a single cell. But instead you&apos;re assigned to a shuffle shard of 4 name servers out of a fleet of thousands. A targeted DDoS attack against your specific hosted zone can only affect the small subset of customers who also happen to share servers with you (and even then, only partially). It&apos;s isolation via statistical independence.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;real-world-implementations&quot;&gt;Real-World Implementations&lt;/h2&gt;
&lt;h3 id=&quot;slack-az-siloing&quot;&gt;Slack: AZ Siloing&lt;/h3&gt;
&lt;p&gt;After the June 2021 incident (&lt;a href=&quot;https://slack.engineering/slacks-migration-to-a-cellular-architecture/&quot;&gt;blog article&lt;/a&gt;), at Slack we redesigned our architecture around availability zone boundaries. Each AZ became a cell and traffic is routed into an AZ only by an Envoy based layer. Services communicate strictly within their AZ. If an AZ becomes unhealthy, operators can &quot;drain&quot; it in under 5 minutes, then redirecting traffic at the edge while allowing the in-flight requests to complete gracefully.&lt;/p&gt;
&lt;p&gt;The key innovation here is the whole &lt;strong&gt;zero external dependencies during evacuation&lt;/strong&gt;. The draining mechanism cannot rely on any infrastructure in the AZ being drained as that AZ might be completely unreachable. The evacuation thus has to be driven from outside the failure boundary.&lt;/p&gt;
&lt;h3 id=&quot;salesforce-hyperforce-cells&quot;&gt;Salesforce: Hyperforce Cells&lt;/h3&gt;
&lt;p&gt;Salesforce&apos;s Hyperforce is perhaps the most ambitious cellular deployment in the industry. At first it seemed unecessary to me. But over the years as I saw one incidents after other and how they were all contained - the setup grew on me. Each cell is a software defined construct managing hundreds of thousands of customers. Cells are deployed across at least 3 availability zones, with immutable infrastructure and zero-trust security (every internal API call is authenticated, no implicit access).&lt;/p&gt;
&lt;p&gt;Hyperforce processes over 100 billion requests per day. When we hit AWS&apos;s hard limit of 250,000 IP addresses per VPC (a constraint of Network Address Usage quotas), we ended up collaborating with AWS to implement prefix delegation. This ended up extending their capacity to 1 million IPs. The lesson: even with perfect cellular isolation, you can still hit scaling cliffs at the infrastructure layer.&lt;/p&gt;
&lt;h3 id=&quot;aws-physalia-millions-of-tiny-databases&quot;&gt;AWS Physalia: Millions of Tiny Databases&lt;/h3&gt;
&lt;p&gt;The most extreme example of cell-based design is AWS Physalia, the configuration master for Elastic Block Store (EBS). Each EBS volume in an availability zone gets its own dedicated cell which is a 7-node Paxos cluster spread across different racks and power supplies. Physalia is not a one big database managing all EBS metadata; it&apos;s &lt;strong&gt;millions of tiny databases&lt;/strong&gt;, each responsible for a single volume. The blast radius of a Physalia node failure is just one server rack. Not one AZ or one region, but just one rack. The grain of isolation is the individual resource making this cellular design taken to its logical extreme.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;building-your-first-cell&quot;&gt;Building Your First Cell&lt;/h2&gt;
&lt;p&gt;If you&apos;re convinced this is the right pattern for your scale and risk profile, here&apos;s the high-level playbook:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Choose Your Partition Key&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is the most important decision you can make. The key must align with the grain of your service. The goal is to not have small number of very large partitions, and if you end up having one - there should be a good reason for it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Tenant ID&lt;/strong&gt;: For B2B SaaS (Slack, Salesforce) where customers expect isolation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;User ID&lt;/strong&gt;: For B2C apps where you want noisy neighbor mitigation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resource ID&lt;/strong&gt;: For storage services (S3, EBS) where data locality matters&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Geography&lt;/strong&gt;: For regulatory compliance (GDPR, data residency)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;2. Design for Static Stability&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Your cell router must work even if your control plane is completely down let alone degraded. Cache the routing map locally. Use DNS, not database lookups, for cell discovery if possible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Deploy in Waves&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Never roll out changes to all cells simultaneously. Use a canary cell with low-risk, low-traffic first. Monitor for say 30-60 minutes, and then deploy to 10% of cells, then 25%, then 100%. If anything looks off via error rates, latency, memory usage - HALT!. Pause the rollout and investigate. Automated rollback is non-negotiable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Embrace Asynchronous Everything&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Cross cell communication must be asynchronous. This is non-negotiable for isolation. Use event buses (like Kafka, Kinesis, EventBridge) for tenant-to-tenant interactions. Stream data to a global analytics warehouse for cross-cell reporting. Keep the hot path isolated across cells.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5. Test the Blast Radius&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Someone, at some time, for some reason will break the rule. You just need to catch it in time and stop the practice from propagating. Run chaos experiments. Deliberately kill a cell in production (with approvals, obviously). Verify that only 1/N of users are affected and that evacuation happens automatically. If the failure spreads beyond the cell boundary, your isolation isn&apos;t real, and your next few weeks are blocked.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-bottom-line&quot;&gt;The Bottom Line&lt;/h2&gt;
&lt;p&gt;Cell-based architecture is not a best practice. It&apos;s a &lt;strong&gt;scaling strategy for when the cost of downtime exceeds the cost of operational complexity&lt;/strong&gt;. If you&apos;re serving millions of users and a mere 5 minute outage ends up costing you six figures in revenue and customer trust, cells are your insurance policy. If you&apos;re a small team building an MVP, they&apos;re probably overkill (but keep the pattern in your back pocket for when you graduate). The trade-off is stark: you&apos;re duplicating infrastructure, increasing your monthly cloud bill, and adding layers of orchestration complexity. In return, you get mathematical predictability. When something fails - and it eventually for sure will - you know exactly how many users are affected: &lt;code&gt;1/N&lt;/code&gt;. That&apos;s my friend is not an estimate but it&apos;s a bare minimum guarantee.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Resilience at scale isn&apos;t about preventing failure. It&apos;s about making failure boring.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you&apos;re aspiring to five-nines availability (99.999% uptime, which translates to just 5 minutes of downtime per year), cells aren&apos;t optional but foundation. Not because they prevent failure, but because they make failure survivable. And in a distributed system, that&apos;s the only guarantee worth having.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Serverless Tax: Why Your LLM Wrapper is Bleeding You Dry]]></title><description><![CDATA[You shipped the shiny new "AI Powered" feature. Your users are happy. ChatGPT is answering questions, Claude is summarizing documents, and…]]></description><link>https://mayankraj.com/blog/serverless-tax-llm-lambda-economics</link><guid isPermaLink="false">https://mayankraj.com/blog/serverless-tax-llm-lambda-economics</guid><pubDate>Mon, 07 Jul 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;You shipped the shiny new &quot;AI Powered&quot; feature. Your users are happy. ChatGPT is answering questions, Claude is summarizing documents, and your product feels like magic - you can finally adda a new subscription tier. Then the AWS bill arrives and it&apos;s triple what you expected. You check the Lambda metrics and nothing looks unusual. The functions are fast, the memory usage is reasonable. But the charges keep climbing, and you can&apos;t figure out why.&lt;/p&gt;
&lt;p&gt;Now imagine this - You hired a plumber to fix your clogged sink. They show up, do their checks and diagnose the issue in 90 seconds. They then sit in your kitchen for the next 45 minutes waiting for a specialty part to arrive from the warehouse. This whole time, their $200/hour meter is running. You&apos;re paying them to make small talks, sip water and scroll. This is what happens when you wrap an LLM in a Lambda function. The function does maybe 5 milliseconds of actual work of parsing the request, calling the API, formatting the response. Then it sits there, idle, while GPT-12.9-mini-high thinks for 20 seconds and streams back tokens at a leisurely 100 per second. AWS doesn&apos;t care that your function is waiting. The meter is running.&lt;/p&gt;
&lt;p&gt;This for you - is the Serverless Tax.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-gb-second-reckoning&quot;&gt;The GB-Second Reckoning&lt;/h2&gt;
&lt;p&gt;Let&apos;s talk about how Lambda actually bills you as it reveals the entire problem. AWS charges you for two things: the number of requests (a flat $0.20 per million) and the duration of those requests, measured in gigabyte-seconds.&lt;/p&gt;
&lt;p&gt;A gigabyte-second (GB-s) is a product of allocated memory of the function (in gigabytes) and how long it runs (in seconds). A function with 1 GB of memory running for 10 seconds consumes 10 GB-s. Simple enough. The current rate is about $0.0000166667 per GB-second for x86 instances (Arm Graviton2 is about 20% cheaper, but let&apos;s keep the math simple).&lt;/p&gt;
&lt;p&gt;Nowhere in here did we talk about utilization. That&apos;s left on you. So if you use 1% or 90% of CPU - you&apos;re charged the same. CPU power is not a separate dimension. When you allocate memory to a Lambda function, you signal AWS to proportionally allocates vCPU. Want a full vCPU? Allocate 1,769 MB of memory. Want more compute? Crank the memory higher. This design made perfect sense in the era of compute heavy workloads like image processing, data transformation, video transcoding. More memory meant more CPU, and you paid for what you used.&lt;/p&gt;
&lt;p&gt;But LLMs broke the model.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-53x-problem&quot;&gt;The 53x Problem&lt;/h2&gt;
&lt;p&gt;Let me show you a typical Retrieval-Augmented Generation (RAG) request. Your user asks a question. Your Lambda function:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Receives the HTTP request (2ms of CPU time)&lt;/li&gt;
&lt;li&gt;Queries your vector database for relevant context (50ms)&lt;/li&gt;
&lt;li&gt;Constructs the prompt with retrieved chunks (3ms)&lt;/li&gt;
&lt;li&gt;Calls the OpenAI API with &lt;code&gt;stream: true&lt;/code&gt; (0.5ms to fire the request)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Waits 4,800ms while GPT-4 generates 500 tokens and streams them back&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Forwards the stream to the user (negligible CPU, it&apos;s just proxying bytes)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Total &lt;em&gt;actual compute&lt;/em&gt;: ~55 milliseconds. Total &lt;em&gt;billed duration&lt;/em&gt;: 4,855 milliseconds. You just paid for &lt;strong&gt;88 times more time&lt;/strong&gt; than you actually &quot;used&quot; the CPU.&lt;/p&gt;
&lt;p&gt;(And this assumes a fast model. If you&apos;re using a slower model for complex reasoning, or Claude in extended &quot;thinking&quot; mode, you might be waiting 30 seconds or more. The ratio gets worse.)&lt;/p&gt;
&lt;p&gt;Here&apos;s the math for a single request with a 1 GB function:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Billed duration: 4.855 seconds&lt;/li&gt;
&lt;li&gt;Memory allocation: 1 GB&lt;/li&gt;
&lt;li&gt;Cost: 4.855 × 0.0000166667 = &lt;strong&gt;$0.0000809 per request&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That doesn&apos;t sound like much until you&apos;re handling 10 million requests per month. Then it&apos;s $809 in Lambda costs alone. The funny part is that the actual LLM tokens might cost you $300. &lt;strong&gt;Your wrapper function is more expensive than the AI.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;When your infrastructure bill exceeds your API bill, the architecture is upside-down.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;response-streaming-a-beautiful-lie&quot;&gt;Response Streaming: A Beautiful Lie&lt;/h2&gt;
&lt;p&gt;AWS Lambda&apos;s &lt;code&gt;streamifyResponse&lt;/code&gt; decorator was positioned to solve this problem. In theory instead of buffering the entire LLM response in memory (which could hit the 6 MB payload limit for traditional Lambda responses), you can stream the response back directly to the client. The user sees tokens appear in near real time.&lt;/p&gt;
&lt;p&gt;Turns out that streaming doesn&apos;t reduce billed duration. The function is still well and alive, connected for the entire duration of the stream. If anything, streaming makes the problem &lt;em&gt;worse&lt;/em&gt;, because now you&apos;re optimizing for user experience at the expense of infrastructure efficiency. The function can&apos;t terminate early. It has to sit there, forwarding bytes, until the last token arrives. The friction was removed which gave developers a cheat code to push the changes out. And we all know what happens with those &quot;TODO: Replace with optimized....&quot; comments.&lt;/p&gt;
&lt;p&gt;Plus if your response exceeds 6 MB (maybe you&apos;re returning a long-form reasoning trace or a large document), you then start paying egress fees at the rate of $0.008 per GB beyond the first 6 MB. For most chat applications, this is irrelevant but for others dealing with images or files - it&apos;s a new cost.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;(To be clear: streaming is the right architectural choice for UX in my opinion. But it doesn&apos;t fix the economic problem. It just makes it more palatable to the end user while making your CFO nervous.)&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-solutions---good-old-web-sockets&quot;&gt;The Solutions - Good Old Web Sockets&lt;/h2&gt;
&lt;p&gt;Alright, enough doom. Let&apos;s talk about how to fix this. If your application is conversational (a chatbot, a coding assistant, a real-time agent), WebSockets flip the model entirely. Instead of the Lambda staying alive for the duration of the LLM call, you keep a &lt;em&gt;persistent connection&lt;/em&gt; between the client and API Gateway. Lambda functions are invoked &lt;em&gt;only when messages are sent or received&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The economics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;WebSocket connection time: $0.25 per million &lt;em&gt;minutes&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;Lambda invocation: Only when there&apos;s actual work to do&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A user might stay connected for 20 minutes but only send 5 messages. You&apos;re paying connection fees (pennies) and 5 Lambda invocations (also pennies). You&apos;re not paying for the 19 minutes of idle time between messages.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;# Lambda is invoked ONLY on message receipt
def websocket_handler(event, context):
    connection_id = event[&apos;requestContext&apos;][&apos;connectionId&apos;]
    user_message = json.loads(event[&apos;body&apos;])[&apos;message&apos;]

    # Call LLM (async, with callback to API Gateway)
    llm_client.generate_streaming(
        prompt=user_message,
        callback=lambda token: send_to_websocket(connection_id, token)
    )

    return {&quot;statusCode&quot;: 200}  # Lambda terminates; connection stays open w/t API Gateway
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The streaming tokens get pushed to the client via the WebSocket, but the Lambda isn&apos;t sitting there playing the relay. It fires off the request and initiates a friendship. The connection is maintained by API Gateway, which costs almost nothing.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-async-native-future&quot;&gt;The Async-Native Future&lt;/h2&gt;
&lt;p&gt;The fundamental mismatch is that serverless was designed for a &lt;em&gt;compute-heavy&lt;/em&gt; workloads whereas LLMs are &lt;em&gt;I/O-heavy&lt;/em&gt;. It&apos;s the wrong tool for the job and that shows up in observability gaps, pricing, scaling etc.&lt;/p&gt;
&lt;p&gt;We&apos;ll need serious retooling for these use cases. We&apos;ve not had a demand for such I/O heavy jobs, at the scale we&apos;re currently seeing. Cloud platforms will need to bill for &lt;em&gt;active CPU cycles&lt;/em&gt;, not wall-clock time. We&apos;re seeing early experiments and implementations on this already..&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The secret:&lt;/strong&gt; Serverless works beautifully when you&apos;re computing. It bleeds you dry when you&apos;re waiting. LLMs force you to wait.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Matryoshka Architecture: Why Envelope Encryption Isn't About Security (It's About Economics)]]></title><description><![CDATA[Imagine a locked dropbox at the supermarket next to you. You've receive a letter but you don't get it right away. Instead of handing it over…]]></description><link>https://mayankraj.com/blog/envelope-encryption-matryoshka-architecture</link><guid isPermaLink="false">https://mayankraj.com/blog/envelope-encryption-matryoshka-architecture</guid><pubDate>Fri, 06 Jun 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Imagine a locked dropbox at the supermarket next to you. You&apos;ve receive a letter but you don&apos;t get it right away. Instead of handing it over to you immediately, the staff hands you this funky looking key card. That key card doesn&apos;t open the dropbox directly but instead it grants you access to a vault at the far end of the counter. Here lies the &lt;em&gt;real&lt;/em&gt; key to the dropbox. Once you retrieve that key, you walk back to the dropbox, unlock it, read your letter, and then immediately destroy the key. The key card remains safe in the vault, ready to issue new keys for the next letter.&lt;/p&gt;
&lt;p&gt;This is envelope encryption in play ...plus with literal envelopes. And if your first reaction is &quot;that sounds unnecessarily complicated,&quot; you&apos;re absolutely right. But here&apos;s the detail: when you&apos;re protecting at scale of petabytes, processing 1M transactions per second, and your cloud bill is approaching a few million dollars per month, &lt;em&gt;unnecessary complexity becomes economic necessity&lt;/em&gt;. The goal of &quot;encrypt everything&quot; isn&apos;t theoretical anymore but instead it&apos;s a regulatory mandate enforced by PCI DSS, HIPAA, GDPR and your CEO&apos;s grandma. The question isn&apos;t whether to encrypt. It&apos;s how to encrypt. Doing that without bankrupting the company or grinding your database to a halt.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-problem-when-encrypt-everything-meets-reality&quot;&gt;The Problem: When &quot;Encrypt Everything&quot; Meets Reality&lt;/h2&gt;
&lt;p&gt;Let&apos;s say you&apos;re an engineer at a SaaS company. Your CISO walks into a meeting and announces the new policy: all data at rest must be encrypted, effective immediately. You respond with - &quot;EOD, Done&quot;. Sounds simple enough. You&apos;ve got a master key already in AWS Key Management Service. You ask claude to create aPR that would encrypt every database record before it hits disk.&lt;/p&gt;
&lt;p&gt;The PR goes to review, and you get your first review - &quot;This is in critical path, let&apos;s test for regression and benchmark this&quot;. You run the first benchmark. Low and behold your transaction throughput drops by 80%. Your P99 latency balloons from 50ms to 2,000ms. Your AWS bill now includes a line item labeled &quot;KMS API Calls&quot; with a projected cost of $1,000,000 per month. &lt;em&gt;What the hell just happened?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;You&apos;ve hit three fundamental failure modes of direct encryption at scale:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Network Bottleneck&lt;/strong&gt;: Your master key lives in a remote KMS as storing it on the application server itself would defeat the purpose of using a dedicated KMS. Every encryption operation requires a costly network round-trip, even if it&apos;s within region, or availability zone. At a speedy 300ms per call, encrypting 1,000 database records sequentially takes 300 &lt;em&gt;seconds&lt;/em&gt;!!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Cryptographic Wear-out&lt;/strong&gt;: AES-GCM, the industry-standard encryption algorithm has a hard limit in it&apos;s maths: you can&apos;t use the same key for more than 2³² messages without risking IV collisions. In a high-throughput system, 2³² messages might occur in hours if not minutes.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Blast Radius&lt;/strong&gt;: A single master key protecting your entire database is a single point of failure. If that key is compromised through a memory exploit, an insider threat, or a simple misconfigured IAM policy then every byte of data you&apos;ve ever encrypted is now exposed.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;blockquote&gt;
&lt;p&gt;Direct encryption doesn&apos;t scale. Not because the cryptography is weak, but because of everything surrounding it. Theres the physics of networks and the mathematics of ciphers imposed constraints that no amount of hardware can overcome.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-envelope-pattern-keys-protecting-keys&quot;&gt;The Envelope Pattern: Keys Protecting Keys&lt;/h2&gt;
&lt;p&gt;Envelope encryption takes a stab at this problems by introducing a two-tiered hierarchy:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Data Encryption Key (DEK)&lt;/strong&gt;: A short-lived symmetric AES-256 key that is actually the one encrypting your data. This key is unique per object i.e. one key per database record, one key per file, one key per user session. It lives in plaintext only in memory, only for the duration of the encryption operation, and is immediately destroyed afterward.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Encryption Key (KEK)&lt;/strong&gt;: A long-lived master key that lives in the KMS and never leaves it in plaintext, with the sole job of encrypting (wrap) and decrypting (unwrap) DEKs. The KEK is stored in a Hardware Security Module (HSM) which are tamper-resistant cryptographic processor that will self-destruct its keys if someone tries to physically breach it.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This dance of encryption looks something like this...&lt;/p&gt;
&lt;h3 id=&quot;encryption-wrapping-the-envelope&quot;&gt;Encryption (Wrapping the Envelope)&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Generate&lt;/strong&gt;: Your application asks the KMS for a new data key. The KMS generates a 256-bit AES key and returns two versions: a plaintext DEK (for immediate use) and a wrapped DEK (encrypted with the KEK).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Encrypt Locally&lt;/strong&gt;: Your application uses the plaintext DEK to encrypt the data. This happens in-memory, at microsecond speeds, often using hardware-accelerated AES instructions (Intel AES-NI).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bundle&lt;/strong&gt;: The application stores the encrypted data (ciphertext) alongside the wrapped DEK. Think of it as a sealed envelope with the key taped to the outside but the key itself is locked in a safe.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Destroy&lt;/strong&gt;: The application overwrites the plaintext DEK in memory. It&apos;s gone. If an attacker scrapes your application&apos;s memory five seconds later, they&apos;ll find nothing in that memory space.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;decryption-opening-the-envelope&quot;&gt;Decryption (Opening the Envelope)&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Retrieve&lt;/strong&gt;: The application reads the ciphertext and the wrapped DEK alongside it from storage. Remember this DEK is specific to this ciphertext only.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unwrap&lt;/strong&gt;: The application sends the wrapped DEK to the KMS. The KMS verifies the requester&apos;s identity (say via IAM roles), decrypts the wrapped DEK using the KEK inside the HSM and then returns the plaintext DEK.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decrypt Locally&lt;/strong&gt;: The application uses the plaintext DEK to decrypt the data locally. This again happens locally at microsecond speeds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Destroy&lt;/strong&gt;: The plaintext DEK is purged from memory.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;It&apos;s important to note the pattern that &lt;strong&gt;you only call the KMS once per session or per object and not once per byte&lt;/strong&gt;. After unwrapping the DEK for that object, all the subsequent cryptographic operations happen locally. The 300ms network penalty is paid once, not a million times. You&apos;ve also reduced your blast radius in the process, with DEK only responsible for a small number of objects.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-three-wins-performance-security-and-economics&quot;&gt;The Three Wins: Performance, Security, and Economics&lt;/h2&gt;
&lt;h3 id=&quot;win-1-performance-the-network-bottleneck-solved&quot;&gt;Win #1: Performance (The Network Bottleneck Solved)&lt;/h3&gt;
&lt;p&gt;By encrypting data locally with the DEK, envelope encryption eliminates any need of sending bulk data over the network for encryption. A 1 GB file encrypted with AES-256 on a modern CPU takes less than a second. The same file sent to a remote KMS for encryption would take minutes ...that is if it&apos;s not first blocked by KMS API&apos;s 4 KB payload limit. Benchmarks from PostgreSQL show that envelope encryption introduces ~5–15% overhead on transactional throughput. For Redis this overhead is less than ~3%.&lt;/p&gt;
&lt;h3 id=&quot;win-2-security-the-blast-radius-contained&quot;&gt;Win #2: Security (The Blast Radius Contained)&lt;/h3&gt;
&lt;p&gt;Because each object has its own unique DEK, any compromise is localized. Let&apos;s say an attacker exploits a vulnerability in your application and manages to extract a DEK from memory. They&apos;ve gained access to exactly &lt;em&gt;one&lt;/em&gt; database record which is the one currently being processed. The KEK remains safely in the HSM, and the millions of other DEKs remain wrapped and unreadable. This is the &quot;defense in depth&quot; strategy applied to cryptography where a single breach doesn&apos;t cascade into a total system compromise.&lt;/p&gt;
&lt;h3 id=&quot;win-3-economics-the-atlassian-story&quot;&gt;Win #3: Economics (The Atlassian Story)&lt;/h3&gt;
&lt;p&gt;Managed KMS services charge by API call&apos;s. AWS KMS for example charges $0.03 per 10,000 requests. If you&apos;re calling the KMS for every database write in a high-traffic application then those pennies add up &lt;em&gt;fast&lt;/em&gt;. Engineers at Atlassian calculated that their direct-encryption model would have cost approximately &lt;strong&gt;$1,000,000 per month&lt;/strong&gt; in KMS fees alone. By implementing envelope encryption with a local cryptographic materials cache, they reduced this to &lt;strong&gt;$2,500 per month&lt;/strong&gt;. That&apos;s a 400x cost reduction!&lt;/p&gt;
&lt;p&gt;Let me repeat that: it&apos;s the same security posture, the same compliance checkboxes, but &lt;strong&gt;$997,500 less per month&lt;/strong&gt;. Envelope encryption isn&apos;t just a technical pattern at this point but it&apos;s a business survival strategy.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-operational-superpowers&quot;&gt;The Operational Superpowers&lt;/h2&gt;
&lt;p&gt;Beyond the core performance and security wins, envelope encryption unlocks two operational capabilities that are impossible with direct encryption:&lt;/p&gt;
&lt;h3 id=&quot;superpower-1-shallow-key-rotation&quot;&gt;Superpower #1: Shallow Key Rotation&lt;/h3&gt;
&lt;p&gt;Regulatory frameworks like PCI DSS mandate that encryption keys be rotated periodically to protect from any breach&apos;s impact. In a direct encryption world, rotating the master key means decrypting and re-encrypting every byte of data. For a petabyte-scale database this is a multi-week high-risk operation.&lt;/p&gt;
&lt;p&gt;With envelope encryption the key rotation is trivial. You rotate the KEK in the KMS and then you re-wrap the DEKs with the underlying data ciphertext never changing. Because DEKs are only often 32 bytes, re-wrapping millions of them can happen as a background process or lazily on the next access. No downtime. No bulk data migration. Just a metadata update for DEK mappings.&lt;/p&gt;
&lt;h3 id=&quot;superpower-2-cryptographic-shredding&quot;&gt;Superpower #2: Cryptographic Shredding&lt;/h3&gt;
&lt;p&gt;The GDPR&apos;s &quot;Right to Erasure&quot; mandate that organizations be able to prove data has been permanently deleted. In a distributed cloud environment it nearly impossible to ensure data is wiped from all backups, replicas, and logs is notoriously difficult.&lt;/p&gt;
&lt;p&gt;Envelope encryption provides an elegant solution: delete the key and not the data. If you destroy the DEK (or the KEK that wraps it), the underlying ciphertext get demoted to garbage bytes. The data still exists on disk somewhere but it&apos;s indistinguishable from random noise. This for us is &quot;cryptographic shredding,&quot; and it provides a tamper-proof audit trail for regulators.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-hardware-root-of-trust-why-hsms-matter&quot;&gt;The Hardware Root of Trust: Why HSMs Matter&lt;/h2&gt;
&lt;p&gt;The integrity of the entire envelope encryption system hinges on one assumption which is that the KEK is secure. If the KEK itself is compromised, the entire hierarchy collapses. This is why we generally use Hardware Security Module (HSM) for KEK&apos;s. An HSM is a dedicated cryptographic processor with actual physical tamper resistance. The KEK takes birth, lives and dies within the wall of HSMs i.e. it never exits the device in plaintext. All wrapping and unwrapping operations occur inside the HSM&apos;s hardened circuitry. If an attacker attempts to physically breach the device, the HSM detects the intrusion and automatically &quot;zeroizes&quot; its memory. This erases the keys before they can be extracted. For enterprise-grade systems or mission critical data, the HSM must be validated to FIPS 140-3 Level 3. It guarantees strong tamper resistance and identity-based authentication. This is the &quot;root of trust&quot; that makes the entire envelope pattern viable.&lt;/p&gt;
&lt;p&gt;HSMs are fascinating and while at Salesforce I had the opportunity to work on the HSM layer that powered the whole of Salesforce ecosystem. I hope to write more about HSMs, and how critical they are to the whole &quot;trust us with our data&quot; claims someday.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-not-to-use-envelope-encryption&quot;&gt;When NOT to Use Envelope Encryption&lt;/h2&gt;
&lt;p&gt;Okay let&apos;s talk about the anti-pattern. Envelope encryption is not a one size fit&apos;s all, and introducing it prematurely can create unnecessary complexity.&lt;/p&gt;
&lt;p&gt;Don&apos;t use envelope encryption if...&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Your dataset is small (under 1 GB) and accessed infrequently&lt;/li&gt;
&lt;li&gt;You&apos;re a startup with 10 users and no compliance requirements (you have bigger problems - users!)&lt;/li&gt;
&lt;li&gt;Your application is CPU-bound, not I/O-bound (the overhead might matter more than the benefit)&lt;/li&gt;
&lt;li&gt;You&apos;re using full-disk encryption (which already provides data-at-rest protection)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Envelope encryption shines at scale. If you&apos;re processing millions of transactions every day, storing sensitive data across the globe, or operating in a regulated industry, it&apos;s the only pattern that reconciles the competing demands of security, performance, and cost.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-post-quantum-future&quot;&gt;The Post-Quantum Future&lt;/h2&gt;
&lt;p&gt;In this series I&apos;ve hyped up every cryptographic approach only to say - the cryptographic landscape is threatened by Quantum progress. A &quot;cryptographically relevant quantum computer&quot; (CRQC) would break the RSA and ECC algorithms which today protect many KEKs. This is the &quot;harvest now, decrypt later&quot; threat wherein the attackers storing encrypted data today with the intent to decrypt it once quantum computers are viable.&lt;/p&gt;
&lt;p&gt;Envelope encryption provides some level of cryptographic agility. Because the KEK is separated from the DEK, we can upgrade KEKs to post-quantum cryptography (PQC) algorithms (eg lattice-based cryptography) without needing to re-encrypt their massive underlying data sets. As long as the AES-256 DEKs remain secure (symmetric encryption is largely quantum resistant with sufficient key lengths), the data remains protected.&lt;/p&gt;
&lt;p&gt;This is the biggest draw of envelope encryption: it&apos;s not just about solving today&apos;s problems. It&apos;s about building a system that can evolve as the threat landscape changes.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-reveal-its-not-about-security&quot;&gt;The Reveal: It&apos;s Not About Security&lt;/h2&gt;
&lt;p&gt;Int he end, envelope encryption doesn&apos;t make your cryptography stronger. AES-256 is already unbreakable with todays tech. The cipher itself doesn&apos;t care whether you&apos;re using a single key or a million keys. What envelope encryption does is make encryption scalable, operable, and economically viable at the petabyte scale. It transforms encryption from a theoretical best practice into a production ready system that can sustain 100,000 transactions per second, all without melting your infrastructure or your budget.&lt;/p&gt;
&lt;p&gt;The Matryoshka doll isn&apos;t just a clever metaphor (thank you fi you thought so). It&apos;s a fundamental architectural pattern that separates the concerns of data transformation (fast, local, ephemeral) from key governance (secure, centralized, long-lived). Envelope encryption is the reconciliation of cryptographic theory with operational reality. It&apos;s the difference between &quot;encrypt everything&quot; as a mandate and &quot;encrypt everything&quot; as a sustainable practice.&lt;/p&gt;
&lt;p&gt;And to answer the question we started with - if you&apos;re an engineer at a SaaS company staring at a million dollar KMS bill, it&apos;s also the difference between a successful security program and a resume generating event.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Math AWS Doesn't Want You to Do: Why S3 Intelligent Tiering Might Be Costing You More]]></title><description><![CDATA[You read the blog, validated the theory via ChatGPT and now you've enabled S3 Intelligent Tiering on your production bucket. The AWS…]]></description><link>https://mayankraj.com/blog/s3-intelligent-tiering-hidden-costs-math</link><guid isPermaLink="false">https://mayankraj.com/blog/s3-intelligent-tiering-hidden-costs-math</guid><pubDate>Wed, 21 May 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;You read the blog, validated the theory via ChatGPT and now you&apos;ve enabled S3 Intelligent Tiering on your production bucket. The AWS marketing promised &quot;automatic cost optimization&quot; and &quot;no operational overhead.&quot; Dream. Set it and forget it, they said. A month later, your bill arrives. You are excited to boast the % reduction in your next 1:1.&lt;/p&gt;
&lt;p&gt;And ...It&apos;s higher than before.&lt;/p&gt;
&lt;p&gt;Not by a little but by thousands of dollars. You check CloudWatch and verified that the data migrated correctly, you even open a support ticket. Everything is &quot;working as designed.&quot; That&apos;s when you realize: S3 Intelligent Tiering isn&apos;t broken. &lt;strong&gt;You just didn&apos;t do the math right.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-storage-unit-that-charges-per-box&quot;&gt;The Storage Unit That Charges Per Box&lt;/h2&gt;
&lt;p&gt;Imagine renting a storage unit for your stuff. The storage company offers you a deal: &quot;We&apos;ll automatically move your boxes to cheaper warehouses when you don&apos;t access them! You&apos;ll save money!&quot; Sounds perfect. Here&apos;s the catch: They charge you $5 per month for every box, to monitor which boxes you&apos;re accessing and then move them around. If you have 10 large boxes with expensive furniture, that&apos;s $50/month in monitoring fees to save $200/month in warehouse rent. Great deal.&lt;/p&gt;
&lt;p&gt;But what if you have 1,000 small boxes? Tiny boxes. Say shoe boxes filled with old receipts, envelopes with birthday cards. Now you&apos;re paying $5,000/month in monitoring fees to save maybe $1,000/month in warehouse rent. The automation service costs magnitudes more than the storage it&apos;s supposed to optimize. You&apos;d have been better off leaving everything in the expensive warehouse and canceling the &quot;smart&quot; service.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;This is S3 Intelligent Tiering for small objects.&lt;/strong&gt; And most people don&apos;t realize it until the bill arrives - very much guilty of this myself.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;how-s3-intelligent-tiering-actually-works&quot;&gt;How S3 Intelligent Tiering Actually Works&lt;/h2&gt;
&lt;p&gt;Let&apos;s dismantle the mechanism so that it&apos;s easy to spot the patterns where it cracks...&lt;/p&gt;
&lt;p&gt;S3 Intelligent Tiering is a storage class that monitors access patterns on individual objects and automatically moves them between tiers:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Frequent Access Tier&lt;/strong&gt; (default): Same price as S3 Standard (~$0.023/GB). Millisecond latency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Infrequent Access Tier&lt;/strong&gt; (30 days idle): 40% cheaper (~$0.0125/GB). Still millisecond latency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Archive Instant Access Tier&lt;/strong&gt; (90 days idle): 68% cheaper. Still instant access.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Archive Access&lt;/strong&gt; (optional, 90 days idle): 71% cheaper, but 3-5 hour retrieval.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deep Archive Access&lt;/strong&gt; (optional, 180 days idle): 95% cheaper, but 12 hour retrieval.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The idea by itself is elegant: Hot data stays fast and accessible and cold data gets cheap. If you suddenly try to access archived data, S3 will automatically promote it back to the Frequent Access tier (for free i.e no retrieval fee). You never have to think about tiers again.&lt;/p&gt;
&lt;p&gt;The automation fairy handles everything while you sleep. &lt;strong&gt;...Or does it?&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-00025-tax-you-didnt-notice&quot;&gt;The $0.0025 Tax You Didn&apos;t Notice&lt;/h2&gt;
&lt;p&gt;Here&apos;s the line item AWS doesn&apos;t bold in the pricing page:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Monitoring and automation charge: $0.0025 per 1,000 objects per month&lt;/strong&gt; (for objects ≥128 KB).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Two and a half &lt;em&gt;thousandths&lt;/em&gt; of a cent per object. That&apos;s... nothing, right? Who cares about fractions of a penny? Let&apos;s do the math...&lt;/p&gt;
&lt;h3 id=&quot;scenario-1-large-number-of-large-objects&quot;&gt;Scenario 1: Large Number of Large Objects&lt;/h3&gt;
&lt;p&gt;You have &lt;strong&gt;1 TB of data&lt;/strong&gt; stored as &lt;strong&gt;10 MB video files&lt;/strong&gt;. That&apos;s roughly 100,000 objects.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Monitoring fee&lt;/strong&gt;: 100,000 / 1,000 × $0.0025 = &lt;strong&gt;$0.25/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storage savings&lt;/strong&gt; (moving to Infrequent Access): 1,000 GB × ($0.023 - $0.0125) = &lt;strong&gt;$10.50/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Net savings&lt;/strong&gt;: $10.50 - $0.25 = &lt;strong&gt;$10.25/month&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Beautiful. This is what Intelligent Tiering was designed for and this is the use case AWS shows you in the case studies.&lt;/p&gt;
&lt;h3 id=&quot;scenario-2-small-number-of-small-objects&quot;&gt;Scenario 2: Small number of Small Objects&lt;/h3&gt;
&lt;p&gt;You have &lt;strong&gt;1 TB of data&lt;/strong&gt; stored as &lt;strong&gt;256 KB JSON logs&lt;/strong&gt;. That&apos;s roughly 4,000,000 objects.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Monitoring fee&lt;/strong&gt;: 4,000,000 / 1,000 × $0.0025 = &lt;strong&gt;$10.00/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storage savings&lt;/strong&gt; (moving to Infrequent Access): 1,000 GB × ($0.023 - $0.0125) = &lt;strong&gt;$10.50/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Net savings&lt;/strong&gt;: $10.50 - $10.00 = &lt;strong&gt;$0.50/month&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You&apos;re paying $10/month to save $10.50/month. Your ROI just evaporated. And this assumes &lt;em&gt;100% of your data reaches the Infrequent Access tier&lt;/em&gt;, which it won&apos;t (more on that later).&lt;/p&gt;
&lt;h3 id=&quot;scenario-3-large-number-of-small-objects&quot;&gt;Scenario 3: Large number of Small Objects&lt;/h3&gt;
&lt;p&gt;You now have &lt;strong&gt;19 TB of data&lt;/strong&gt; stored as &lt;strong&gt;300 KB files&lt;/strong&gt;. That&apos;s roughly 66,000,000 objects. For logs, per-minute data dumps in any large scale, or high throughput scenarios, or even services wherein there&apos;s user submitted artefacts (eg images on Instagram) - this is very common.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Monitoring fee&lt;/strong&gt;: 66,000,000 / 1,000 × $0.0025 = &lt;strong&gt;$165/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Storage savings&lt;/strong&gt;: 19,000 GB × ($0.025 - $0.0125) = &lt;strong&gt;$237.50/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Net savings&lt;/strong&gt;: $237.50 - $165.00 = &lt;strong&gt;$72.50/month&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You just burned $165/month on monitoring fees just to save $237.50/month in storage. That&apos;s a 70% reduction in your expected savings. And we &lt;em&gt;still&lt;/em&gt; haven&apos;t accounted for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The transition fee to move the data into Intelligent Tiering ($660 one-time for 66M objects)&lt;/li&gt;
&lt;li&gt;The request fees for your application&apos;s API calls&lt;/li&gt;
&lt;li&gt;The KMS encryption overhead (if you&apos;re not using S3 Bucket Keys)&lt;/li&gt;
&lt;li&gt;The metadata storage tax in archive tiers&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the real case, the bill was &lt;strong&gt;$2,000/month&lt;/strong&gt; instead of the expected $400. The monitoring fee was just the beginning of the cost cascade.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-128-kb-trap-where-small-objects-go-to-die&quot;&gt;The 128 KB Trap: Where Small Objects Go to Die&lt;/h2&gt;
&lt;p&gt;In 2021, AWS updated S3 Intelligent Tiering with a &quot;cost-saving&quot; feature: &lt;strong&gt;Objects smaller than 128 KB are exempt from the monitoring fee&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Sounds generous ...which is very unlike AWS! AWS is giving small objects a free pass, right?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wrong.&lt;/strong&gt; Here&apos;s what they don&apos;t tell you: If an object is under 128 KB, it&apos;s exempt from the monitoring fee &lt;em&gt;because it&apos;s also exempt from tiering!!&lt;/em&gt;. These objects &lt;strong&gt;never move&lt;/strong&gt; to cheaper tiers. They stay in the Frequent Access tier forever, billed at S3 Standard rates.&lt;/p&gt;
&lt;p&gt;Let&apos;s say you have a dataset with:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;50 million objects under 128 KB (microservice logs, metadata files)&lt;/li&gt;
&lt;li&gt;10 million objects over 128 KB (some larger files)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You enable Intelligent Tiering expecting the whole dataset to benefit. What actually happens:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The 50M small objects stay in Frequent Access ($0.023/GB), paying S3 Standard prices forever&lt;/li&gt;
&lt;li&gt;The 10M larger objects get monitored (costing you $25/month) and &lt;em&gt;maybe&lt;/em&gt; move to cheaper tiers&lt;/li&gt;
&lt;li&gt;You now have the complexity of a &quot;managed storage class&quot; with almost none of the benefits&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The worst part is that you&apos;re under a false sense of security that your bucket needs no further work!! For high cardinality datasets with lots of tiny files, S3 Intelligent Tiering is &lt;strong&gt;strictly worse&lt;/strong&gt; than S3 Standard. At best you get the same cost but with a more complicated storage architecture and even less predictability.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-break-even-formula-do-this-before-you-enable&quot;&gt;The Break-Even Formula (Do This Before You Enable)&lt;/h2&gt;
&lt;p&gt;Let&apos;s derive the actual math. When is S3 Intelligent Tiering profitable?&lt;/p&gt;
&lt;p&gt;For sake of simplicity, let&apos;s say that the monitoring fee ($0.0025 per 1,000 objects) must be less than the storage savings (difference between S3 Standard and Infrequent Access tier):&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Monitoring Cost &amp;#x3C; Storage Savings
(N / 1000) × $0.0025 &amp;#x3C; G × ($0.023 - $0.0125)

Where:
  N = Number of objects ≥128 KB
  G = Total storage in GB
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Simplifying (and assuming 100% of data reaches Infrequent Access):&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;N × $0.0000025 &amp;#x3C; G × $0.0125
N / G &amp;#x3C; $0.0125 / $0.0000025
N / G &amp;#x3C; 5,000 objects per GB
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Since &lt;code&gt;N / G&lt;/code&gt; is the inverse of average object size in GB:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Average Object Size &gt; 1 GB / 5,000
Average Object Size &gt; 0.0002 GB
Average Object Size &gt; ~204.8 KB
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Translation:&lt;/strong&gt; If your average object size is less than ~205 KB then the S3 Intelligent Tiering will cost more than S3 Standard, even if 100% of your data is cold. That&apos;s the best case scenario. In practice, not all data reaches the Infrequent tier (some objects get accessed before the 30-day threshold). So the real break-even point is closer to &lt;strong&gt;300-500 KB average object size&lt;/strong&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you don&apos;t know your average object size, you&apos;re gambling with your cloud bill.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-hidden-costs-they-dont-mention&quot;&gt;The Hidden Costs They Don&apos;t Mention&lt;/h2&gt;
&lt;p&gt;The monitoring fee is just the most visible tax. Here are the other costs that make S3 Intelligent Tiering a minefield.&lt;/p&gt;
&lt;h3 id=&quot;1-the-transition-fee-the-3-million-entry-toll&quot;&gt;1. The Transition Fee: The $3 Million Entry Toll&lt;/h3&gt;
&lt;p&gt;AWS charges &lt;strong&gt;$0.01 per 1,000 objects&lt;/strong&gt; to transition data from S3 Standard to Intelligent Tiering via lifecycle policies.&lt;/p&gt;
&lt;p&gt;For the Canva engineering team (which reportedly has 300 billion objects) transitioning to Intelligent Tiering would cost &lt;strong&gt;$3,000,000 upfront&lt;/strong&gt;. Even for a modest bucket with 100 million objects of any size, that&apos;s a $1,000 entry fee. If your monthly savings is only $100, you&apos;ll need 10 months just to break even on the transition cost.&lt;/p&gt;
&lt;p&gt;You can bypass this fee but only for future objects. You can upload new data directly to Intelligent Tiering (using &lt;code&gt;x-amz-storage-class&lt;/code&gt; header). But this requires code changes, which contradicts the &quot;zero operational overhead&quot; promise. And it means all new data starts in the Frequent Access tier, accruing monitoring fees immediately.&lt;/p&gt;
&lt;h3 id=&quot;2-the-40-kb-metadata-tax&quot;&gt;2. The 40 KB Metadata Tax&lt;/h3&gt;
&lt;p&gt;When objects move to the Archive Access or Deep Archive Access tiers, AWS charges you for &lt;strong&gt;40 KB of metadata overhead per object&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;8 KB billed at S3 Standard rates (for object name and attributes)&lt;/li&gt;
&lt;li&gt;32 KB billed at Glacier rates (for indexing)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For a 1 MB file, 40 KB is negligible (4% overhead). But for a 128 KB file, it&apos;s a &lt;strong&gt;31% increase&lt;/strong&gt; in billable storage. And if you&apos;re archiving billions of tiny objects? The metadata can cost more than the actual data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Example:&lt;/strong&gt; Archiving 1 billion objects of 1 KB each to Glacier Deep Archive.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data storage&lt;/strong&gt;: 1,024 GB × $0.00099/GB = &lt;strong&gt;$1.01/month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metadata overhead&lt;/strong&gt;: 40 KB × 1B objects = 38.1 TB × $0.00099/GB = &lt;strong&gt;$39.60/month&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Well the metadata now costs 40× more than the data. This is why the &quot;Average Object Size&quot; metric isn&apos;t just important but it&apos;s existential.&lt;/p&gt;
&lt;h3 id=&quot;3-the-promotion-reset-loop&quot;&gt;3. The Promotion Reset Loop&lt;/h3&gt;
&lt;p&gt;When you access an object in the Infrequent or Archive tiers, it signals the Intelligent Tiering to automatically promotes it back to the Frequent Access tier. Free promotion, no retrieval fee! Sounds great ...but here&apos;s the problem: &lt;strong&gt;The 30-day idle clock resets.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If you have a compliance scanner, a backup verification tool, or an audit system that touches every object once every 29 days, your data will never make the move to a cheaper tier. You&apos;ll pay the monitoring fee every month, but 100% of your data will stay in Frequent Access (billed at S3 Standard rates). You&apos;re paying for automation that can&apos;t activate because your access pattern is just frequent enough to keep resetting the timer.&lt;/p&gt;
&lt;h3 id=&quot;4-the-one-way-door-problem&quot;&gt;4. The &quot;One-Way Door&quot; Problem&lt;/h3&gt;
&lt;p&gt;Once you transition data into S3 Intelligent Tiering via a lifecycle policy, there&apos;s no automated way to move it back to S3 Standard. That is - If you discover six months later that the monitoring fees are eating your savings then you&apos;re stuck. Moving data out of Intelligent Tiering requires:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A manual operation (e.g., S3 Batch Operations)&lt;/li&gt;
&lt;li&gt;Paying &lt;code&gt;PUT&lt;/code&gt; request fees ($0.005 per 1,000 objects)&lt;/li&gt;
&lt;li&gt;Dealing with potential data transfer costs&lt;/li&gt;
&lt;li&gt;Rewriting lifecycle policies&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is the &quot;set it and forget it&quot; trap: You set it, you forget it, and then you&apos;re financially locked in.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;when-s3-intelligent-tiering-actually-works&quot;&gt;When S3 Intelligent Tiering Actually Works&lt;/h2&gt;
&lt;p&gt;Let me be very clear about this - S3 Intelligent Tiering isn&apos;t a scam but it&apos;s just wildly misunderstood. There are workloads where it&apos;s genuinely the best choice and I almost always default to it. Just any other tool - you have to use the right one for the right job.&lt;/p&gt;
&lt;h3 id=&quot;the-sweet-spot-large-objects-chaotic-access&quot;&gt;The Sweet Spot: Large Objects, Chaotic Access&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;I would consider Intelligent Tiering when:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Average object size &gt; 500 KB&lt;/strong&gt; (preferably &gt; 1 MB)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unpredictable access patterns&lt;/strong&gt; (you genuinely don&apos;t know which files will be accessed when)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Long retention periods&lt;/strong&gt; (data lives for months and years, not days or weeks)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;High-frequency spot access&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Can use SSE-KMS&lt;/strong&gt; (Without S3 Bucket Keys enabled, AWS KMS fees will destroy you)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If my use case doesn&apos;t fit this - I would need some good convincing to look at Intelligent Tiering. Lifecycle rules may be the best option for transitions in such cases - ones that are tuned for the specific needs.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-real-set-it-and-forget-it-lie&quot;&gt;The Real &quot;Set It and Forget It&quot; Lie&lt;/h2&gt;
&lt;p&gt;The central issue with S3 Intelligent Tiering isn&apos;t technical but instead it&apos;s epistemological. AWS marketing materials presents it as a revolutionary &quot;hands-off&quot; solution, which trains engineers to treat it as a configuration checkbox rather than a proper decision requiring ongoing validation.&lt;/p&gt;
&lt;p&gt;In reality, S3 cost optimization is not an event but it&apos;s an ongoing process:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Audit&lt;/strong&gt; your object size distribution (~quarterly)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model&lt;/strong&gt; the costs (monitoring fee + metadata + transitions vs. storage savings)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test&lt;/strong&gt; on a single representative prefix before enabling globally&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Monitor&lt;/strong&gt; the actual costs vs. projections (~monthly)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Refine&lt;/strong&gt; by disabling on prefixes where ROI is negative (~quarterly)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is the opposite of &quot;set it and forget it.&quot; It&apos;s &quot;set it, measure it, course-correct, and continuously optimize.&quot; Build automations, add it to ceremonies and most importantly - document the reasoning so that the next person can follow through.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;takeaways&quot;&gt;Takeaways&lt;/h2&gt;
&lt;p&gt;If you&apos;re walking away from this article thinking &quot;S3 Intelligent Tiering is bad,&quot; you missed the point. Let me reframe:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;S3 Intelligent Tiering is excellent for:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Large objects (≥1 MB average)&lt;/li&gt;
&lt;li&gt;Unpredictable, bursty access patterns&lt;/li&gt;
&lt;li&gt;Long-lived datasets with low object cardinality&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;S3 Intelligent Tiering is catastrophic for:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Small objects (&amp;#x3C;200 KB average)&lt;/li&gt;
&lt;li&gt;Predictable &quot;hot then cold&quot; lifecycles&lt;/li&gt;
&lt;li&gt;High-cardinality datasets (billions of files)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The tool isn&apos;t the problem. The &quot;set it and forget it&quot; framing is the problem. AWS gave you a sophisticated tool but they marketed it as a magic wand. The difference matters. And that could either make your next review cycle a breeze or a difficult conversation.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: Elliptic Curve Cryptography (ECC) — The Efficient Enigma]]></title><description><![CDATA[Why modern privacy relies on playing billiards on a mathematical doughnut. If you have worked in software long enough - you have interacted…]]></description><link>https://mayankraj.com/blog/iipt-pt8-elliptic-curve-cryptography</link><guid isPermaLink="false">https://mayankraj.com/blog/iipt-pt8-elliptic-curve-cryptography</guid><pubDate>Tue, 11 Mar 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Why modern privacy relies on playing billiards on a mathematical doughnut.&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;If you have worked in software long enough - you have interacted with an RSA certificate. You generate a 2048-bit key, paste it somewhere, and move on with your life. The math is complex, the keys are long. If you&apos;ve been paying attention, you might have noticed GitHub defaulting away from RSA to smaller, faster Ed25519 keys for SSH authentication. Then you hear about ECC and wonder why anyone bothered with anything else.&lt;/p&gt;
&lt;p&gt;Think of cryptography like transporting a secret across enemy territory. RSA&apos;s approach is to hide a needle in an enormous haystack—make the haystack big enough, and nobody can find your needle. ECC takes a completely different path: imagine a game of billiards where you hit a ball, it bounces around the table millions of times, and even if I tell you the start and end positions, you &lt;em&gt;cannot&lt;/em&gt; figure out how many times it bounced. That one-way hardness is what now secures your WhatsApp messages, your Bitcoin wallet, and increasingly, every HTTPS connection you make.&lt;/p&gt;
&lt;p&gt;In this deep dive, we strip away the algebraic geometry to reveal the simple mechanics underneath. We&apos;ll explore why &quot;adding&quot; points on a curve is the secret to smaller keys, how the &quot;Pac-Man&quot; effect of finite fields turns smooth curves into secure scatter plots, and why the NSA might have backdoored the very standards we use today &lt;strong&gt;(yes, really)&lt;/strong&gt;. No conspiracy theories here - I promise.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-heavyweight-vs-the-gymnast&quot;&gt;The Heavyweight vs. The Gymnast&lt;/h2&gt;
&lt;p&gt;In a previous installment, we dissected RSA, which in every sense the granddaddy of public-key cryptography. RSA&apos;s security rests on a very simple bet: multiplying two massive prime numbers is trivial, but factoring the result back into those primes is computationally brutal.&lt;/p&gt;
&lt;p&gt;The problem? RSA is struggling to keep up.&lt;/p&gt;
&lt;p&gt;To match the security of a 256-bit AES key (the gold standard for symmetric encryption), RSA would need a key that&apos;s over &lt;strong&gt;15,360 bits long&lt;/strong&gt;. That&apos;s like needing a dump truck to carry a single envelope. Your mobile device is shuffling around these monstrous numbers, burning CPU cycles and battery life - all to &quot;securely&quot; view the next reel. The math works. &lt;strong&gt;The efficiency doesn&apos;t.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Enter ECC, the shiny new gymnast to RSA&apos;s heavyweight.&lt;/p&gt;
&lt;p&gt;ECC offers military-grade security with keys that fit on a sticky note. A &lt;strong&gt;256-bit ECC key&lt;/strong&gt; provides roughly the same security as that 15,360-bit RSA behemoth. All by abandoning the &quot;needle in a haystack&quot; approach entirely and exploiting a different mathematical labyrinth.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ECC doesn&apos;t make the haystack bigger but instead it makes the search algorithm itself useless.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;playing-billiards-on-a-curve&quot;&gt;Playing Billiards on a Curve&lt;/h2&gt;
&lt;p&gt;At the heart of ECC lies an elegant piece of mathematics: the &lt;strong&gt;elliptic curve&lt;/strong&gt;. The name may fool you but these curves have nothing to do with ellipses. The name is a leftover from their use in calculating the arc length of ellipses centuries ago ...and mathematicians not stressing about naming conventions like you and me do.&lt;/p&gt;
&lt;h3 id=&quot;the-equation&quot;&gt;The Equation&lt;/h3&gt;
&lt;p&gt;An elliptic curve is defined by the &lt;strong&gt;Weierstrass equation&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;$$y^2 = x^3 + ax + b$$&lt;/p&gt;
&lt;p&gt;Where $a$ and $b$ are constants that determine the curve&apos;s shape. There&apos;s one constraint: the curve must be &quot;smooth&quot; (no cusps or self-intersections), which requires $4a^3 + 27b^2 \neq 0$. Plot this equation, and you get a distinctive shape—imagine a sideways bell curve with a smooth hump. It&apos;s symmetric about the x-axis, which becomes crucial for our &quot;game.&quot;&lt;/p&gt;
&lt;p&gt;[DIAGRAM SUGGESTION: Show a smooth elliptic curve with labeled axes and the characteristic &quot;hump&quot; shape]&lt;/p&gt;
&lt;h3 id=&quot;the-game-point-addition&quot;&gt;The Game: Point Addition&lt;/h3&gt;
&lt;p&gt;Let&apos;s look under the hood now. We&apos;ll start by defining a strange operation called &quot;point addition&quot;. It lets us combine two points on the curve to get a third point.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Rule:&lt;/strong&gt; Draw a straight line through any two points ($P$ and $Q$) on an elliptic curve. That line will &lt;em&gt;always&lt;/em&gt; intersect the curve at exactly one more point.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Move:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Draw a line through points $P$ and $Q$&lt;/li&gt;
&lt;li&gt;Find where the line hits the curve again (call this $R&apos;$)&lt;/li&gt;
&lt;li&gt;Reflect $R&apos;$ across the x-axis to get $R$&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;We define: $P + Q = R$&lt;/p&gt;
&lt;p&gt;That&apos;s it!! That&apos;s all that there is to elliptic curve &quot;addition.&quot; It doesn&apos;t look like normal addition, and it shouldn&apos;t, as the curve is playing by its own rules here. The curve doesn&apos;t care about your intuition, &lt;strong&gt;It has its own physics.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;[DIAGRAM SUGGESTION: Animated or step-by-step visual showing P + Q = R with the reflection]&lt;/p&gt;
&lt;h3 id=&quot;point-doubling-when-p-meets-p&quot;&gt;Point Doubling: When P Meets P&lt;/h3&gt;
&lt;p&gt;What happens when $P$ and $Q$ are the same point? You can&apos;t draw a unique line through a single point... or can you?&lt;/p&gt;
&lt;p&gt;When doubling a point, we use the &lt;strong&gt;tangent line&lt;/strong&gt; at that point. The tangent touches the curve exactly at $P$. By the same rules it will intersect the curve at exactly one other point. Reflect that point, and you have $2P$. The curve is stubborn: even when you give it the same point twice, it still produces a meaningful answer.&lt;/p&gt;
&lt;p&gt;This operation, &lt;strong&gt;point doubling&lt;/strong&gt; is the key to making ECC computationally efficient.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-trapdoor-where-one-way-becomes-no-way&quot;&gt;The Trapdoor: Where One-Way Becomes No-Way&lt;/h2&gt;
&lt;p&gt;Now we can define the actual hard problem that secures ECC.&lt;/p&gt;
&lt;h3 id=&quot;the-setup&quot;&gt;The Setup&lt;/h3&gt;
&lt;p&gt;Every ECC system starts with a &lt;strong&gt;Generator Point&lt;/strong&gt; $G$. It is a specific point on the curve that everyone agrees on. Think of it as the starting position of our billiard ball.&lt;/p&gt;
&lt;h3 id=&quot;the-easy-way-scalar-multiplication&quot;&gt;The Easy Way: Scalar Multiplication&lt;/h3&gt;
&lt;p&gt;We can &quot;multiply&quot; a point by a number through repeated addition:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$2G = G + G$ (point doubling)&lt;/li&gt;
&lt;li&gt;$3G = 2G + G$&lt;/li&gt;
&lt;li&gt;$kG = G + G + G + ...$ ($k$ times)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&apos;s the beautiful part: we actually don&apos;t need to add $G$ to itself $k$ times. Using &lt;strong&gt;&quot;Double and Add&quot;&lt;/strong&gt;, we can compute $kG$ in roughly $\log_2(k)$ operations. This algorithm is greedy and it takes shortcuts wherever it can.&lt;/p&gt;
&lt;p&gt;For example, to compute $100G$:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$100 = 64 + 32 + 4$ (in binary: 1100100)&lt;/li&gt;
&lt;li&gt;Calculate $2G, 4G, 8G, 16G, 32G, 64G$ through successive doubling&lt;/li&gt;
&lt;li&gt;Add the relevant results: $64G + 32G + 4G = 100G$&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This, my friends was 100 additions in about 7 operations. Scale this up: we can compute $k \cdot G$ for astronomically large $k$ (think 256-bit numbers) in mere milliseconds. The CPU is &lt;em&gt;happy&lt;/em&gt; doing this.&lt;/p&gt;
&lt;h3 id=&quot;the-hard-way-the-discrete-logarithm-problem&quot;&gt;The Hard Way: The Discrete Logarithm Problem&lt;/h3&gt;
&lt;p&gt;Now flip the problem:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Given the starting point $G$ and the ending point $Q = kG$, find $k$.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is the &lt;strong&gt;Elliptic Curve Discrete Logarithm Problem (ECDLP)&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;We know where the ball started and also where it landed. &lt;strong&gt;But how many times did it bounce?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;With properly chosen curves and sufficiently large numbers, the best known algorithms to solve this problem would take longer than the age of the universe. The math is elegant. The computational wall is absolute.&lt;/p&gt;
&lt;h3 id=&quot;your-keys&quot;&gt;Your Keys&lt;/h3&gt;
&lt;p&gt;This asymmetry gives us our key pair:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Private Key:&lt;/strong&gt; The scalar $k$ (a random number you keep secret)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Public Key:&lt;/strong&gt; The point $Q = kG$ (you share this with the world)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Computing $Q$ from $k$ is trivial—your phone does it in milliseconds. Computing $k$ from $Q$? The sun will burn out first.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;finite-fields-when-smooth-curves-become-chaos&quot;&gt;Finite Fields: When Smooth Curves Become Chaos&lt;/h2&gt;
&lt;p&gt;There&apos;s a problem with everything we&apos;ve discussed: computers hate infinity.&lt;/p&gt;
&lt;p&gt;The smooth, continuous curves we&apos;ve been visualizing have infinitely many points with infinite-precision decimal coordinates. Your CPU can&apos;t work with that &lt;strong&gt;(and neither can your RAM ...or your patience)&lt;/strong&gt;. We need integers.&lt;/p&gt;
&lt;h3 id=&quot;the-pac-man-effect&quot;&gt;The Pac-Man Effect&lt;/h3&gt;
&lt;p&gt;To get around this, we restrict our curve to be a &lt;strong&gt;finite field&lt;/strong&gt;. Instead of working with all real numbers, we work with integers modulo a large prime number $p$.&lt;/p&gt;
&lt;p&gt;Our curve equation then becomes:
$$y^2 \equiv x^3 + ax + b \pmod{p}$$&lt;/p&gt;
&lt;p&gt;Every calculation simply wraps around when it exceeds $p$.&lt;/p&gt;
&lt;p&gt;Next up: imagine our billiard table has Pac-Man physics. When the ball reaches the right edge, it teleports to the left. When it goes off the top, it reappears at the bottom. Now imagine playing billiards on this table. After millions of bounces, could you trace the path backward just by knowing the start and end positions?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nope.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The modular arithmetic creates this &quot;teleportation&quot; effect. Addition and multiplication still work, but the results wrap around unpredictably. The mathematical relationships that make the curve secure are preserved, but the visual patterns are obliterated.&lt;/p&gt;
&lt;h3 id=&quot;the-transformation&quot;&gt;The Transformation&lt;/h3&gt;
&lt;p&gt;Something strange happens when we apply this constraint:&lt;/p&gt;
&lt;p&gt;The smooth, elegant curve &lt;strong&gt;vanishes&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In its place: a seemingly random cloud of scattered points. No pattern. No shape. Just dots sprayed across a grid. The curve hasn&apos;t changed mathematically.&lt;/p&gt;
&lt;p&gt;[DIAGRAM SUGGESTION: Side-by-side comparison of smooth curve vs. finite field scatter plot]&lt;/p&gt;
&lt;p&gt;This is the genius of finite field ECC: all the algebraic properties we need for cryptography survive, while the &quot;structure&quot; an attacker might exploit is annihilated. The math works perfectly but the pattern recognition fails completely.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-efficiency-showdown-why-ecc-won&quot;&gt;The Efficiency Showdown: Why ECC Won&lt;/h2&gt;
&lt;p&gt;Let&apos;s talk numbers. Here&apos;s why the entire internet is migrating to ECC:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Security Level (bits)&lt;/th&gt;
&lt;th&gt;RSA Key Size&lt;/th&gt;
&lt;th&gt;ECC Key Size&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;1,024&lt;/td&gt;
&lt;td&gt;160&lt;/td&gt;
&lt;td&gt;6.4x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;112&lt;/td&gt;
&lt;td&gt;2,048&lt;/td&gt;
&lt;td&gt;224&lt;/td&gt;
&lt;td&gt;9.1x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;3,072&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;12x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;192&lt;/td&gt;
&lt;td&gt;7,680&lt;/td&gt;
&lt;td&gt;384&lt;/td&gt;
&lt;td&gt;20x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;15,360&lt;/td&gt;
&lt;td&gt;512&lt;/td&gt;
&lt;td&gt;30x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The ratio gets &lt;em&gt;worse&lt;/em&gt; for RSA as security requirements increase. The more security you need, the more ECC outperforms RSA.&lt;/p&gt;
&lt;h3 id=&quot;the-real-world-impact&quot;&gt;The Real-World Impact&lt;/h3&gt;
&lt;p&gt;A 256-bit ECC key vs. a 3,072-bit RSA key (both providing 128-bit security):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Key Generation:&lt;/strong&gt; ECC keys are generated ~10x faster&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Signatures:&lt;/strong&gt; ECDSA is ~10x faster than RSA signatures (equivalent security)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bandwidth:&lt;/strong&gt; ECC public keys are 12x smaller&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Battery:&lt;/strong&gt; Mobile devices consume significantly less power &lt;strong&gt;(your phone thanks you)&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is why TLS 1.3 defaults to ECDHE (Elliptic Curve Diffie-Hellman Ephemeral) for key exchange. It&apos;s enabled our phone to establish hundreds of secure connections without draining the battery. It&apos;s why Bitcoin can process thousands of transaction signatures every second.&lt;/p&gt;
&lt;p&gt;The efficiency isn&apos;t marginal. It&apos;s transformational.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;implementation-code-you-can-touch&quot;&gt;Implementation: Code You Can Touch&lt;/h2&gt;
&lt;p&gt;Theory is beautiful, but code is truth. Here&apos;s a simplified implementation of point addition on an elliptic curve:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;class Point:
    &quot;&quot;&quot;A point on an elliptic curve y² = x³ + ax + b (mod p)&quot;&quot;&quot;

    def __init__(self, x, y, curve):
        self.x = x
        self.y = y
        self.curve = curve  # Contains a, b, and p

    def __add__(self, other):
        # The identity element: point at infinity
        # (Yes, elliptic curves have their own version of zero)
        if self.x is None:
            return other
        if other.x is None:
            return self

        p = self.curve.p

        # When points are vertical opposites, they cancel out
        # P + (-P) = infinity (the curve shrugs and returns nothing)
        if self.x == other.x and self.y != other.y:
            return Point(None, None, self.curve)

        # Point doubling: P + P
        # The tangent line takes over when we&apos;re adding a point to itself
        if self.x == other.x and self.y == other.y:
            # slope = (3x² + a) / (2y) mod p
            # The curve leans into the tangent here
            numerator = (3 * self.x**2 + self.curve.a) % p
            denominator = (2 * self.y) % p
            slope = (numerator * mod_inverse(denominator, p)) % p
        else:
            # Standard point addition: P + Q
            # Just draw the line and find where it hits
            numerator = (other.y - self.y) % p
            denominator = (other.x - self.x) % p
            slope = (numerator * mod_inverse(denominator, p)) % p

        # Calculate the new point
        # This is where the reflection magic happens
        x3 = (slope**2 - self.x - other.x) % p
        y3 = (slope * (self.x - x3) - self.y) % p

        return Point(x3, y3, self.curve)


def mod_inverse(a, p):
    &quot;&quot;&quot;Fermat&apos;s little theorem saves us from division headaches&quot;&quot;&quot;
    return pow(a, p - 2, p)


def scalar_multiply(k, point):
    &quot;&quot;&quot;
    Double-and-add: the algorithm that makes ECC practical.
    We&apos;re computing k * Point without actually adding k times.
    &quot;&quot;&quot;
    result = Point(None, None, point.curve)  # Start at infinity
    addend = point

    while k:
        if k &amp;#x26; 1:  # If lowest bit is set, add current point
            result = result + addend
        addend = addend + addend  # Double (always)
        k &gt;&gt;= 1  # Shift right, on to the next bit

    return result
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This code demonstrates the core mechanics:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Point addition&lt;/strong&gt; with slope calculation (the geometry becomes algebra)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The modulo operator&lt;/strong&gt; wrapping all values (Pac-Man physics)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Double-and-add&lt;/strong&gt; for efficient scalar multiplication (the shortcut that makes it all practical)&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;production-code&quot;&gt;Production Code&lt;/h3&gt;
&lt;p&gt;For anything real, use a library. Here&apos;s what that looks like:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;from cryptography.hazmat.primitives.asymmetric import ec
from cryptography.hazmat.backends import default_backend

# Generate a key pair on P-256
# (The library handles all the scary math)
private_key = ec.generate_private_key(ec.SECP256R1(), default_backend())
public_key = private_key.public_key()

# For Curve25519 key exchange (the cool kids&apos; choice)
from cryptography.hazmat.primitives.asymmetric.x25519 import X25519PrivateKey

private_key = X25519PrivateKey.generate()
public_key = private_key.public_key()
# Done. That&apos;s it. The library did the heavy lifting.
&lt;/code&gt;&lt;/pre&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-quantum-horizon-eccs-expiration-date&quot;&gt;The Quantum Horizon: ECC&apos;s Expiration Date&lt;/h2&gt;
&lt;p&gt;Here&apos;s the uncomfortable truth that keeps cryptographers up at night: ECC has an expiration date on the box.&lt;/p&gt;
&lt;h3 id=&quot;the-threat&quot;&gt;The Threat&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Shor&apos;s Algorithm&lt;/strong&gt;, discovered in 1994, can solve both the integer factorization problem (breaking RSA) and the discrete logarithm problem (breaking ECC). All in polynomial time on a quantum computer.&lt;/p&gt;
&lt;p&gt;And here&apos;s the cruel irony: ECC&apos;s efficiency today that make it compelling becomes its downfall. Because ECC uses smaller keys, a quantum computer needs &lt;em&gt;fewer qubits&lt;/em&gt; to crack it.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Algorithm&lt;/th&gt;
&lt;th&gt;Key Size&lt;/th&gt;
&lt;th&gt;Qubits Needed (estimate)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RSA-2048&lt;/td&gt;
&lt;td&gt;2,048&lt;/td&gt;
&lt;td&gt;~4,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECC-256&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;~2,330&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id=&quot;the-timeline&quot;&gt;The Timeline&lt;/h3&gt;
&lt;p&gt;Current quantum computers have roughly 1,000-1,500 noisy qubits. We&apos;ll need millions of error corrected qubits to run Shor&apos;s Algorithm. By estimates we&apos;re 10-20 years away from cryptographically relevant quantum computers but only time will tell. We&apos;re also a bit distracted with all the LLM stuff today.&lt;/p&gt;
&lt;p&gt;But here&apos;s the catch: &lt;strong&gt;&quot;harvest now, decrypt later.&quot;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Adversaries can record encrypted communications today and decrypt them once quantum computers arrive. If you&apos;re protecting secrets that need to stay secret for decades, say state secrets, medical records, intellectual property then the quantum threat is already here. Your secrets are sitting in someone&apos;s storage, waiting for the future.&lt;/p&gt;
&lt;h3 id=&quot;the-future-post-quantum-cryptography&quot;&gt;The Future: Post-Quantum Cryptography&lt;/h3&gt;
&lt;p&gt;The cryptographic community is already building the next generation of defenses: &lt;strong&gt;Post-Quantum Cryptography (PQC)&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;NIST finalized its first post-quantum standards in 2024:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ML-KEM (Kyber):&lt;/strong&gt; For key encapsulation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ML-DSA (Dilithium):&lt;/strong&gt; For digital signatures&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SLH-DSA (SPHINCS+):&lt;/strong&gt; Hash-based signatures&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These algorithms are based on mathematical problems that even quantum computers can&apos;t solve efficiently like lattice problems, hash functions etc. The trade-off? Larger keys and signatures. We&apos;re going back to the dump truck era, at least partially.&lt;/p&gt;
&lt;h3 id=&quot;when-not-to-use-ecc&quot;&gt;When NOT to Use ECC&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Long-term secrets requiring 30+ year confidentiality:&lt;/strong&gt; Consider hybrid schemes (ECC + post-quantum)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systems that can&apos;t be updated post-deployment:&lt;/strong&gt; Embedded devices, satellites, IoT with 20+ year lifespans&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environments where timing attacks are feasible and you can&apos;t use Curve25519:&lt;/strong&gt; Some older curves require constant-time implementations that are hard to get right&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-takeaway&quot;&gt;The Takeaway&lt;/h2&gt;
&lt;p&gt;Elliptic Curve Cryptography represents the pinnacle of classical cryptography. It&apos;s the most elegant balance of security and efficiency we&apos;ve ever achieved. It transformed mobile security, enabled cryptocurrency, and quietly protects billions of daily transactions.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ECC&apos;s core insight is in the geometric game of bouncing points on a curve, constrained by modular arithmetic, creates a one-way function so powerful that all the computers on Earth couldn&apos;t reverse it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Having said that, nothing lasts forever. As the quantum era approaches, and with it, new mathematical foundations will fundamentally replace the curves we&apos;ve come to trust. But for the next decade or two, ECC remains the workhorse of internet security.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Aurora DSQL vs DynamoDB MRSC: The Global Database Duel]]></title><description><![CDATA[Scene: You're sitting in an architecture review. The CTO leans forward and asks the question that'll decide the next quarter's roadmap: "Can…]]></description><link>https://mayankraj.com/blog/aurora-dsql-vs-dynamodb-mrsc</link><guid isPermaLink="false">https://mayankraj.com/blog/aurora-dsql-vs-dynamodb-mrsc</guid><pubDate>Tue, 04 Feb 2025 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Scene: You&apos;re sitting in an architecture review. The CTO leans forward and asks the question that&apos;ll decide the next quarter&apos;s roadmap: &quot;Can we guarantee that a user in Tokyo sees the exact same inventory count as a user in Frankfurt, instantly, without either of them waiting 200 milliseconds for the other region to phone home?&quot;&lt;/p&gt;
&lt;p&gt;The room goes quiet. Because for two decades, the honest answer was: &quot;Pick two: global, consistent, or fast.&quot;&lt;/p&gt;
&lt;p&gt;Well ! &lt;strong&gt;Not anymore!!&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-city-hall-problem&quot;&gt;The City Hall Problem&lt;/h2&gt;
&lt;p&gt;Let me paint you a picture. Imagine that you&apos;re running a global operation with two very different coordination challenges:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Challenge One: The City Hall Records Office.&lt;/strong&gt; You have birth certificates, property deeds, marriage licenses, business permits and what not. Every document references other documents. A marriage license references two birth certificates. A business permit references a property deed, which references an owner, who has a marriage license, which... you get the idea, right?&lt;/p&gt;
&lt;p&gt;Now zoom out. When someone in the London office updates a property deed, the clerk in New York needs to see that change &lt;em&gt;immediately&lt;/em&gt;. Why? Because they&apos;re about to issue a business permit that depends on it. The system is structured, cross-referenced, and deeply relational. One wrong reference and the entire legal framework collapses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Challenge Two: The FedEx Tracking Network.&lt;/strong&gt; You have millions of packages flying around the planet. Each package has a tracking number, a destination, a weight, and maybe some custom attributes (fragile, perishable, &quot;definitely not a birthday present&quot;). Packages don&apos;t really reference each other.&lt;/p&gt;
&lt;p&gt;When you look up tracking number &lt;code&gt;111222333444555&lt;/code&gt;, you&apos;re only interested in that one package&apos;s status. You don&apos;t care about complex joins across manifests and route tables. You care about speed, scale, and the flexibility to slap new attributes on packages without redesigning the entire tracking database schema.&lt;/p&gt;
&lt;p&gt;This is the duel between &lt;strong&gt;Amazon Aurora DSQL&lt;/strong&gt; and &lt;strong&gt;DynamoDB Multi-Region Strong Consistency (MRSC)&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;City Hall vs. FedEx. Relational rigor vs. NoSQL agility. Both now with a superpower we were told was impossible - &lt;strong&gt;Global Strong Consistency&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;what-changed-spoiler-physics-didnt&quot;&gt;What Changed? (Spoiler: Physics Didn&apos;t)&lt;/h2&gt;
&lt;p&gt;For years, the CAP theorem was gospel. For many it still is. In any distributed system during a network &lt;em&gt;P&lt;/em&gt;artition, you can have &lt;em&gt;C&lt;/em&gt;onsistency &lt;em&gt;or&lt;/em&gt; &lt;em&gt;A&lt;/em&gt;vailability, but not both. Most architects chose their side early:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Team Consistency:&lt;/strong&gt; PostgreSQL, MySQL with synchronous replication. Slow, regional, but correct.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Team Availability:&lt;/strong&gt; Cassandra, original DynamoDB. Fast, global, but eventually consistent.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then Google Spanner came along in 2012 and said, &quot;What if we just... synchronized atomic clocks across data centers?&quot; TrueTime felt like cheating but It was brilliant. Fast forward to 2024, and AWS dropped two services that bring Spanner level strong consistency to the masses...&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Aurora DSQL&lt;/strong&gt; — A distributed SQL database that uses hardware-assisted time synchronization and optimistic concurrency control.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DynamoDB MRSC&lt;/strong&gt; — Multi-region strong consistency for DynamoDB Global Tables using a global journal sequencer.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Both systems achieve &lt;strong&gt;PC/EC&lt;/strong&gt; in the PACELC theorem. Translation: When there&apos;s a partition (P), they choose consistency (C). When everything&apos;s running normally (E), they still choose consistency (C) over latency (L). This is the &quot;hard mode&quot; of distributed databases, and AWS is now offering it with the fan favorite flavour of serverless.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;aurora-dsql-city-hall-goes-serverless&quot;&gt;Aurora DSQL: City Hall Goes Serverless&lt;/h2&gt;
&lt;h3 id=&quot;the-architecture-or-how-to-disaggregate-everything&quot;&gt;The Architecture (Or: How to Disaggregate Everything)&lt;/h3&gt;
&lt;p&gt;With traditional databases a single monolithic process makes up compute, storage, and transaction management. Aurora DSQL saw that, and took a different path...&lt;/p&gt;
&lt;p&gt;Here&apos;s the flow when you run a &lt;code&gt;BEGIN; UPDATE users SET balance = balance - 100 WHERE id = 42; COMMIT;&lt;/code&gt; transaction:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;The Relay&lt;/strong&gt; gets your connection via TLS SNI and routes you to a &lt;strong&gt;Query Processor (QP)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The QP is a stateless PostgreSQL compatible compute node running in a &lt;strong&gt;Firecracker microVM&lt;/strong&gt;. No noisy neighbors. No state. It starts by parsing your SQL, planing your query, and then during the transaction it just &lt;em&gt;remembers&lt;/em&gt; your changes in memory (like a waiter taking your order but not sending it to the kitchen yet).&lt;/li&gt;
&lt;li&gt;When you say &lt;code&gt;COMMIT&lt;/code&gt;, the QP screams the transaction to the &lt;strong&gt;Adjudicator&lt;/strong&gt; — a lease-based coordinator that owns specific key ranges. The Adjudicator uses &lt;strong&gt;Optimistic Concurrency Control (OCC)&lt;/strong&gt;. It checks: &quot;Did anyone else touch these rows since you started?&quot; If not, it proceeds. If yes, the Adjudicator throws up its hands and gives you a PostgreSQL serialization error (&lt;code&gt;40001&lt;/code&gt;), and you must retry.&lt;/li&gt;
&lt;li&gt;If validation passes, the transaction gets committed to the &lt;strong&gt;Journal&lt;/strong&gt; — a distributed log service (think S3-grade durability). The Journal is replicated across two active regions and a witness region. Once the Journal says &quot;committed,&quot; your data is durable.&lt;/li&gt;
&lt;li&gt;Asynchronously, the &lt;strong&gt;Storage Layer&lt;/strong&gt; (independant service) learns about committed transactions from the Journal and updates its local view. When you read, DSQL uses a &quot;pushdown&quot; model: filters and aggregations happen at the storage layer, so only the relevant rows come back to the QP. No shuffling entire table scans across the network.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is how AWS can claim DSQL being &lt;strong&gt;4x faster&lt;/strong&gt; than traditional distributed SQL. Because all of the chatty PostgreSQL protocol round trips (SELECT, UPDATE, UPDATE, UPDATE...) happen locally in the QP&apos;s memory. The QP buffers everything like a patient scribe, taking notes. The global coordination tax is paid exactly once, at &lt;code&gt;COMMIT&lt;/code&gt; time.&lt;/p&gt;
&lt;h3 id=&quot;the-trade-off-optimism-has-a-price&quot;&gt;The Trade-Off: Optimism Has a Price&lt;/h3&gt;
&lt;p&gt;OCC is fantastic when contention is low (different users updating different rows). It&apos;s &lt;em&gt;painful&lt;/em&gt; when contention is high (everyone trying to decrement the same inventory counter simultaneously). In high-contention scenarios, one transaction wins. Everyone else gets a serialization error and retries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Your application must be idempotent.&lt;/strong&gt; Because retrying &quot;charge this credit card&quot; without safeguards is how you get angry customers and chargebacks.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;dynamodb-mrsc-fedex-gets-a-global-sequencer&quot;&gt;DynamoDB MRSC: FedEx Gets a Global Sequencer&lt;/h2&gt;
&lt;h3 id=&quot;the-architecture-or-how-to-make-key-value-coordination-not-terrible&quot;&gt;The Architecture (Or: How to Make Key-Value Coordination Not Terrible)&lt;/h3&gt;
&lt;p&gt;In standard DynamoDB Global Tables each region has its own partition. Writes are local and fast. Replicators then ship changes to other regions asynchronously. If two regions update the same item simultaneously, DynamoDB uses a simple &lt;strong&gt;last-writer-wins (LWW)&lt;/strong&gt; based on timestamps. This is eventual consistency. It works for 90% of use cases (session state, cache invalidation, analytics).&lt;/p&gt;
&lt;p&gt;But what if you&apos;re running a global leaderboard, and two players in different regions both claim to be #1 at the exact same millisecond? LWW picks one based on timestamps, and the other player&apos;s update silently vanishes into the void.&lt;/p&gt;
&lt;p&gt;That&apos;s... not great for player retention (or your app store rating).&lt;/p&gt;
&lt;p&gt;Enter &lt;strong&gt;Multi-Region Strong Consistency (MRSC)&lt;/strong&gt;. Here&apos;s how it works:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;When you write to an MRSC enabled table, the write goes to a &lt;strong&gt;Multi-Region Journal (MRJ)&lt;/strong&gt; first. It assigns a globally unique sequence number to the write.&lt;/li&gt;
&lt;li&gt;The write must be acknowledged by a &lt;strong&gt;quorum of regions&lt;/strong&gt; (say 2 out of 3) before the client gets a success response. This is synchronous replication. Your RPO (Recovery Point Objective) is zero — if Region A explodes, Region B already has your data.&lt;/li&gt;
&lt;li&gt;When you try a strongly consistent read (&lt;code&gt;ConsistentRead=True&lt;/code&gt;), the regional replica sends a &lt;strong&gt;heartbeat&lt;/strong&gt; to other regions or to MRJ to ensure it&apos;s caught up with all globally committed writes. The replica is basically asking, &quot;Hey, did I miss anything?&quot; Only after getting confirmation does it return data.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This ensures you never read stale data. But it also means that a strongly consistent read in MRSC is no longer a local operation. You&apos;re paying cross-region latency for correctness.&lt;/p&gt;
&lt;h3 id=&quot;the-trade-off-consistency-isnt-free-literally&quot;&gt;The Trade-Off: Consistency Isn&apos;t Free (Literally)&lt;/h3&gt;
&lt;p&gt;DynamoDB pricing got a massive overhaul in late 2024: &lt;strong&gt;50% cut&lt;/strong&gt; in on-demand throughput pricing and &lt;strong&gt;67% cut&lt;/strong&gt; in global tables replication costs.&lt;/p&gt;
&lt;p&gt;But here&apos;s the thing: at millions of requests per second, DynamoDB&apos;s per-item pricing model can get expensive. A single RDBMS or DSQL cluster might handle the same query load for less money because you&apos;re paying for compute time, not individual request units. Also, if two regions write to the same item simultaneously, you get a &lt;strong&gt;replicated write conflict exception&lt;/strong&gt;. Just like DSQL&apos;s serialization errors, your application must catch this and retry. Idempotency is non-negotiable.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-performance-showdown&quot;&gt;The Performance Showdown&lt;/h2&gt;
&lt;p&gt;Let&apos;s talk numbers...&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Operation&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Aurora DSQL&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;DynamoDB MRSC&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Local Read&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~1-5ms&lt;/td&gt;
&lt;td&gt;~1-5ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Strong Read (Global)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~1-5ms (snapshot)&lt;/td&gt;
&lt;td&gt;~20-200ms+ (heartbeat)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Write (Interactive)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~1-5ms (buffered)&lt;/td&gt;
&lt;td&gt;~20-200ms+ (synchronous)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Write (Commit/Ack)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2 RTTs (sync)&lt;/td&gt;
&lt;td&gt;1-2 RTTs (sync)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Max Throughput&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Virtually unlimited&lt;/td&gt;
&lt;td&gt;20M+ requests/sec&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The absolute killer feature of DSQL is that &lt;strong&gt;strongly consistent reads are still local&lt;/strong&gt;. Because DSQL uses snapshot isolation tied to physical time, a QP can serve reads from its local storage without consulting other regions.&lt;/p&gt;
&lt;p&gt;DynamoDB MRSC, by contrast, has to send a heartbeat to verify consistency. &lt;strong&gt;Every. Single. Strong. Read.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This is the speed-of-light problem. Tokyo to Frankfurt is roughly 150-200ms round-trip. Just to be on the same page... you can&apos;t cheat physics.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-decision-framework-or-how-to-not-screw-this-up&quot;&gt;The Decision Framework (Or: How to Not Screw This Up)&lt;/h2&gt;
&lt;p&gt;Here&apos;s a mental math I generally use...&lt;/p&gt;
&lt;h3 id=&quot;q-do-i-need-complex-joins-foreign-keys-or-multi-table-transactions&quot;&gt;Q: &quot;Do I need complex joins, foreign keys, or multi-table transactions?&quot;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;If yes:&lt;/strong&gt; Aurora DSQL. DynamoDB doesn&apos;t do joins. You can work around them with Global Secondary Indexes and multiple queries. But that classic case of unnecessary overengineering.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If no:&lt;/strong&gt; Keep reading.&lt;/p&gt;
&lt;h3 id=&quot;q-is-my-schema-stable-or-am-i-still-figuring-out-what-attributes-my-items-need&quot;&gt;Q: &quot;Is my schema stable, or am I still figuring out what attributes my items need?&quot;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;If stable:&lt;/strong&gt; Aurora DSQL is fine. Schema migrations in DSQL are just &lt;code&gt;ALTER TABLE&lt;/code&gt; statements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If volatile:&lt;/strong&gt; DynamoDB. You can add new attributes to items on the fly without &lt;code&gt;ALTER TABLE&lt;/code&gt; hell. Hate to say this in 2025, but this is the superpower of NoSQL.&lt;/p&gt;
&lt;h3 id=&quot;q-whats-my-readwrite-ratio-and-whats-the-access-pattern&quot;&gt;Q: &quot;What&apos;s my read/write ratio, and what&apos;s the access pattern?&quot;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;If mostly reads, complex queries, or ad-hoc analytics:&lt;/strong&gt; Aurora DSQL. SQL is expressive, beautiful and self-documented.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If mostly writes, or simple key lookups at massive scale:&lt;/strong&gt; DynamoDB. It&apos;s built for throughput.&lt;/p&gt;
&lt;h3 id=&quot;q-whats-my-tolerance-for-operational-complexity&quot;&gt;Q: &quot;What&apos;s my tolerance for operational complexity?&quot;&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;If you want maximum abstraction:&lt;/strong&gt; Both are serverless. But DSQL is &lt;em&gt;newer&lt;/em&gt;. It&apos;s in preview. DynamoDB is battle-tested and powers AWS&apos;s own internal systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you&apos;re risk-averse:&lt;/strong&gt; DynamoDB MRSC. It&apos;s a feature flag on an existing, mature service.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-gotchas&quot;&gt;The Gotchas...&lt;/h2&gt;
&lt;h3 id=&quot;for-aurora-dsql&quot;&gt;For Aurora DSQL:&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;No foreign keys (yet).&lt;/strong&gt; You&apos;re left to handle it with application enforced referential integrity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;3,000 row / 10MB limit per transaction.&lt;/strong&gt; If you&apos;re doing bulk imports, break them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preview status means features are missing.&lt;/strong&gt; No PostGIS, no triggers, no full-text search (yet).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sequential primary keys are death.&lt;/strong&gt; Auto-incrementing IDs will create a hot spot where every write hammers the same Adjudicator. Instead use UUIDs or random keys to distribute the load.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;for-dynamodb-mrsc&quot;&gt;For DynamoDB MRSC:&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Strongly consistent reads are slow.&lt;/strong&gt; Plan your access patterns really really really really carefully.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Item size limit is 400KB.&lt;/strong&gt; If you&apos;re storing blobs, use S3 and reference them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Idempotency by design.&lt;/strong&gt; Conflict exceptions require retry logic.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pricing at scale can surprise you.&lt;/strong&gt; Run the cost calculator and compare against DSQL for your specific query profile.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&quot;the-punchline&quot;&gt;The Punchline&lt;/h2&gt;
&lt;p&gt;Just two decades ago if you wanted global strong consistency, you&apos;d built a monolithic Oracle RAC cluster ...prayed to the database gods ..and also accepted that &quot;disaster recovery&quot; meant &quot;restore from backup in 4-6 hours.&quot;&lt;/p&gt;
&lt;p&gt;Today you spin up a serverless DSQL cluster or flip a feature flag on DynamoDB. You get active-active multi-region replication with zero RPO and near-zero RTO! The commit latency is still capped by the speed of light itself, but at least you&apos;re not also fighting decades-old database architectures designed for a pre-cloud world.&lt;/p&gt;
&lt;p&gt;So just to recap - when your CTO asks, &quot;Can we guarantee consistency globally?&quot; the answer is finally, confidently: &lt;strong&gt;&quot;Yes. Now let&apos;s talk about your access patterns.&quot;&lt;/strong&gt;&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: RSA - The Asymmetric Anchor]]></title><description><![CDATA[Fresh from our exploration of Diffie-Hellman's elegant key exchange protocol, we're diving into RSA. It's the algorithm that transformed the…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-rsa-the-asymmetric-anchor</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-rsa-the-asymmetric-anchor</guid><pubDate>Sat, 02 Nov 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Fresh from our exploration of Diffie-Hellman&apos;s elegant key exchange protocol, we&apos;re diving into RSA. It&apos;s the algorithm that transformed the theoretical possibility of public-key cryptography into a practical reality. While DH solved the key exchange puzzle, RSA took it a step ahead and opened up entirely new possibilities with its asymmetric approach to encryption and digital signatures. Today, it stands as a cornerstone of modern cryptography, but faces a formidable challenge. Yes, you guessed it - quantum computing.&lt;/p&gt;
&lt;p&gt;This article builds on the mathematical foundations we&apos;ve explored throughout this series. We&apos;ll revisit everything from the prime numbers that fascinated us in our DES discussion to the modular arithmetic that powered our AES implementations. Now the pieces of the puzzle will start to fit in as we&apos;ll see how these concepts combine in RSA&apos;s ingenious design. While not yet obsolete but RSA&apos;s vulnerability to Shor&apos;s algorithm demands a thorough understanding of its strengths and weaknesses. Grab your favorite beverage as we dissect RSA&apos;s inner workings.&lt;/p&gt;
&lt;h2 id=&quot;the-rsa-story---more-than-just-initials&quot;&gt;The RSA Story - More Than Just Initials&lt;/h2&gt;
&lt;p&gt;The 70s was truely a wild era for cryptography. Computationlly fast symmetric encryption suffered from a killer flaw of key distribution. How do you securely share a secret key with someone across the internet without any guarentee that someone is already intercepting the traffic? You might think - didn&apos;t we discuss Diffie-Hellman&apos;s 1976 masterpiece in the last article? Well yes, and while DH introduced the concept of public-key cryptography but that theory needed a real-world algorithm. And this voice created space for &lt;a href=&quot;https://en.wikipedia.org/wiki/Ron_Rivest&quot;&gt;Rivest&lt;/a&gt;, &lt;a href=&quot;https://en.wikipedia.org/wiki/Adi_Shamir&quot;&gt;Shamir&lt;/a&gt;, and &lt;a href=&quot;https://en.wikipedia.org/wiki/Leonard_Adleman&quot;&gt;Adleman&lt;/a&gt; to rise up. 1997 saw the paper that introduced RSA algorithm. This wasn&apos;t just a paper but instead it was the birth of practical public-key cryptography. Today we can say that it is a foundation for much of our modern digital security infrastructure.&lt;/p&gt;
&lt;p&gt;The genius of RSA? It leverages the computational asymmetry of prime factorization i.e. while multiplying two large primes is easy but factoring their product is incredibly difficult. Even today powerful supercomputers struggle with this task. This asymmetry allows for a public key (used for encryption and verification) and a private key (used for decryption and signing) - the core of public-key cryptography.&lt;/p&gt;
&lt;h2 id=&quot;a-deeper-dive-into-rsas-number-theory-magic&quot;&gt;A Deeper Dive into RSA&apos;s Number Theory Magic&lt;/h2&gt;
&lt;p&gt;RSA&apos;s security isn&apos;t magic but rather elegant mathematics. Let&apos;s dissect the core components and step a bit beyond the surface level:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Prime Numbers: The Foundation of Foundations&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Two incredibly large prime numbers, say &lt;em&gt;p&lt;/em&gt; and &lt;em&gt;q&lt;/em&gt; form the bedrock of RSA. These aren&apos;t just any primes but they&apos;re very large primes, often thousands of bits long. Finding these primes isn&apos;t trivial. We have few tried and tested means to get these two numbers like the Miller-Rabin test. It offer a high probability (but not absolute certainty) that a number is prime. However the probability of error is so incredibly small that for all practical purposes we can consider the numbers we find to be prime. The product of these primes, &lt;em&gt;n = p&lt;/em&gt;q*, is the modulus. Modulus is the public parameter used in both encryption and decryption. The size of &lt;em&gt;n&lt;/em&gt; directly influences the security of the system. A larger *n* implies a much more difficult factoring problem for attackers.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Euler&apos;s Totient Function: Counting the Coprimes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Euler&apos;s totient function, φ(&lt;em&gt;n&lt;/em&gt;), is then used to get co-primes for &lt;em&gt;n&lt;/em&gt;. Co-primes are the positive integers up to &lt;em&gt;n&lt;/em&gt; that are relatively prime to &lt;em&gt;n&lt;/em&gt; i.e. they share no common factors other than 1. For RSA, where &lt;em&gt;n = p&lt;/em&gt;q*, this simplifies beautifully to φ(&lt;em&gt;n&lt;/em&gt;) = (&lt;em&gt;p&lt;/em&gt; - 1)(*q* - 1). While it seems insignificant it really is not. It dictates the structure of the mathematical group within which the encryption and decryption operations take place.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Modular Arithmetic: The Clockwork of Cryptography&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Modular arithmetic is the arithmetic of remainders. We perform operations (addition, multiplication, exponentiation) and then take the remainder after dividing by a modulus (&lt;em&gt;n&lt;/em&gt; in our case). This dance is what&apos;s depicted by the &quot;clock arithmetic&quot; analogy. If you add 5 hours to 10 o&apos;clock on a 12-hour clock, you get 3 o&apos;clock (because 15 mod 12 = 3). This is fundamental to RSA&apos;s efficiency as it helps to keep the numbers within manageable bounds. Just to reiterate we&apos;re dealing with incredibly large numbers to begin with.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Public Exponent (&lt;em&gt;e&lt;/em&gt;): The Public Face&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The public exponent, &lt;em&gt;e&lt;/em&gt;, is a relatively small number (often 65537, which is a Fermat prime) that&apos;s coprime to φ(&lt;em&gt;n&lt;/em&gt;) (their greatest common divisor is 1). This is done to ensures that it has a multiplicative inverse modulo φ(&lt;em&gt;n&lt;/em&gt;). The public key (&lt;em&gt;n&lt;/em&gt;, &lt;em&gt;e&lt;/em&gt;) is openly advertised and freely available to anyone. Somone who wants to send you an encrypted message can do so with this public key. The choice of &lt;em&gt;e&lt;/em&gt; thus involves a tradeoff. A small &lt;em&gt;e&lt;/em&gt; leads to slightly faster encryption but could potentially lead to vulnerabilities. For that reason 65537 is a good balance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Private Exponent (&lt;em&gt;d&lt;/em&gt;): The Secret Keeper&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;We&apos;re now getting into the heart of RSA&apos;s security. The private exponent &lt;em&gt;d&lt;/em&gt; is the multiplicative inverse of &lt;em&gt;e&lt;/em&gt; modulo φ(&lt;em&gt;n&lt;/em&gt;). This means &lt;em&gt;e&lt;/em&gt; * d ≡ 1 (mod φ(&lt;em&gt;n&lt;/em&gt;)). In much simpler terms, if you multiply &lt;em&gt;e&lt;/em&gt; and &lt;em&gt;d&lt;/em&gt; and then take the remainder after dividing by φ(&lt;em&gt;n&lt;/em&gt;), you get 1. Finding &lt;em&gt;d&lt;/em&gt; without knowing &lt;em&gt;p&lt;/em&gt; and &lt;em&gt;q&lt;/em&gt; (and thus φ(&lt;em&gt;n&lt;/em&gt;)) is computationally infeasible using classical algorithms. This is because of the fact that finding this inverse requires knowing the factorization of &lt;em&gt;n&lt;/em&gt;. This is exactly the problem we&apos;re leveraging for security. For all of this to work the private key, (&lt;em&gt;n&lt;/em&gt;, *d*), must remain absolutely secret.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Chinese Remainder Theorem (CRT): The Speed Demon&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The CRT is a mathematical theorem which involves breaking down the problem with large modulus into smaller more manageable subproblems. In the context of RSA decryption it permits performing the decryption operation modulo &lt;em&gt;p&lt;/em&gt; and modulo &lt;em&gt;q&lt;/em&gt; separately. From there we combine the results efficiently to obtain the final result modulo &lt;em&gt;n&lt;/em&gt;. This significantly accelerates decryption often by a factor of four or more. It helps make RSA practical for real-world applications.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;key-generation-the-art-of-crafting-cryptographic-secrets&quot;&gt;Key Generation: The Art of Crafting Cryptographic Secrets&lt;/h2&gt;
&lt;p&gt;The food can only be as good as it&apos;s ingrediants by themself. While RSA relies on strong mathematics but it will all be for nothing of the key pair generation is sloppy. Here&apos;s a detailed breakdown:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Prime Number Generation:&lt;/strong&gt; This is the most computationally expensive part and also the most important. We use some sophisticated probabilistic primality testing algorithms (like the Miller-Rabin test) to generate two large prime numbers. The larger these primes, the more secure the resulting RSA key pair. The process of prime generation deserves some careful setup to ensure that the selected primes do not have any special properties which makes them vulnerable to advanced factoring techniques.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Modulus Calculation:&lt;/strong&gt; The modulus, &lt;em&gt;n&lt;/em&gt;, is calculated as the product of the two primes i.e. &lt;em&gt;n = p&lt;/em&gt;q*. As you can imagine, &lt;em&gt;n&lt;/em&gt; forms the foundation of the public and private keys. The size of *n* (typically expressed in bits) dictates the security level of the RSA system.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Totient Calculation:&lt;/strong&gt; Next, we compute Euler&apos;s totient function, φ(&lt;em&gt;n&lt;/em&gt;) = (&lt;em&gt;p&lt;/em&gt; - 1)(&lt;em&gt;q&lt;/em&gt; - 1). This value is crucial for determining the private exponent. This value, φ(&lt;em&gt;n&lt;/em&gt;), must be kept secret to maintain security of RSA.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Public Exponent Selection:&lt;/strong&gt; We choose the public exponent, &lt;em&gt;e&lt;/em&gt;, such that 1 &amp;#x3C; &lt;em&gt;e&lt;/em&gt; &amp;#x3C; φ(&lt;em&gt;n&lt;/em&gt;) and gcd(&lt;em&gt;e&lt;/em&gt;, φ(&lt;em&gt;n&lt;/em&gt;)) = 1 (i.e., &lt;em&gt;e&lt;/em&gt; and φ(&lt;em&gt;n&lt;/em&gt;) are coprime). A common and generally safe choice is &lt;em&gt;e&lt;/em&gt; = 65537.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Private Exponent Calculation:&lt;/strong&gt; This is where the Extended Euclidean Algorithm joins the party. It&apos;s used to calculate the modular multiplicative inverse of &lt;em&gt;e&lt;/em&gt; modulo φ(&lt;em&gt;n&lt;/em&gt;). This inverse &lt;em&gt;d&lt;/em&gt; is the private exponent. The equation &lt;em&gt;e&lt;/em&gt; * d ≡ 1 (mod φ(&lt;em&gt;n&lt;/em&gt;)) must hold true. The Extended Euclidean Algorithm efficiently finds &lt;em&gt;d&lt;/em&gt;. The calculation of &lt;em&gt;d&lt;/em&gt; requires knowledge of φ(&lt;em&gt;n&lt;/em&gt;), which in turn requires knowledge of &lt;em&gt;p&lt;/em&gt; and &lt;em&gt;q&lt;/em&gt;. It is the secrecy of &lt;em&gt;d&lt;/em&gt; that protects the RSA system at all times. If an attacker could determine *d*, they could decrypt messages intended for the private key holder.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Pair Formation:&lt;/strong&gt; Finally, the public key is formed as the pair (&lt;em&gt;n&lt;/em&gt;, &lt;em&gt;e&lt;/em&gt;), while the private key is (&lt;em&gt;n&lt;/em&gt;, &lt;em&gt;d&lt;/em&gt;). The public key can be freely distributed, while the private key must be kept strictly secret and protected.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This process requires careful attention to detail to ensure the generation of strong and secure RSA key pairs. Any flaws in this generation process can severely compromise the security of the entire system.&lt;/p&gt;
&lt;h2 id=&quot;implementation-details&quot;&gt;Implementation Details&lt;/h2&gt;
&lt;p&gt;The theoretical beauty of RSA needs practical implementation considerations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Key Size:&lt;/strong&gt; Larger keys (2048 bits or more are recommended now) offer stronger security but slower performance. The choice involves a trade-off between security and efficiency.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Padding Schemes:&lt;/strong&gt; Raw RSA is vulnerable. Padding schemes like Optimal Asymmetric Encryption Padding (OAEP) add randomness and structure to the data before encryption. This in turn significantly enhancing security against various attacks (including chosen-ciphertext attacks). PKCS#1 v1.5, while historically prevalent, is now considered insecure.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hardware Acceleration:&lt;/strong&gt; For high-performance applications, hardware-based acceleration via specialized cryptographic processors or Hardware Security Modules (HSMs) is often necessary. This significantly improves performance and security.&lt;/li&gt;
&lt;/ul&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import random
import hashlib
from Crypto.Util.number import getPrime, inverse

def generate_rsa_keys(key_size):
    p = getPrime(key_size//2)
    q = getPrime(key_size//2)
    n = p*q
    phi = (p-1)*(q-1)
    e = 65537                           #Common value
    d = inverse(e, phi)
    return ((n, e), (n, d))

def oaep_padding(message, key_size):
    #Simplified OAEP -  In a real system, use a robust library like pycryptodome
    k = key_size // 8
    mLen = len(message)
    mgf = hashlib.SHA256()
    mgf.update(message)
    maskedDB = mgf.digest()[:k - mLen - 2] + b&apos;\x00&apos; + message
    seed = os.urandom(k//2)
    dbMask = mgf(seed, k - k//2)
    maskedSeed = mgf(maskedDB, k//2) ^ seed
    return maskedSeed + maskedDB
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;security-analysis-rsas-vulnerabilities&quot;&gt;Security Analysis: RSA&apos;s Vulnerabilities&lt;/h2&gt;
&lt;p&gt;RSA&apos;s security while strong for appropriately sized keys and careful implementation is by no means absolute. Key vulnerabilities include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Factoring Attacks:&lt;/strong&gt; RSA&apos;s core relies on the difficulty of factoring large numbers. While currently infeasible for sufficiently large keys with classical computers, quantum computing (via Shor&apos;s algorithm) poses a significant long-term threat.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Side-Channel Attacks:&lt;/strong&gt; These exploit information leaked during computation, such as timing variations (timing attacks) or power consumption patterns (power analysis). Robust countermeasures, including constant-time implementations as well as blinding are essential for mitigating these attacks.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Implementation Errors:&lt;/strong&gt; It&apos;s surprising how many times a poorly written code, inadequate testing, or insecure libraries end up creating vulnerabilities. This is regardless of key size or padding. Using well-vetted libraries, secure coding practices, and rigorous testing are essential for robust security.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Addressing these vulnerabilities requires a multi-pronged approach which involves strong key management, secure implementation, and regular security audits. It&apos;s not just about the math but it&apos;s about secure implementation and proactive risk mitigation.&lt;/p&gt;
&lt;h2 id=&quot;the-looming-threat-of-quantum-computing&quot;&gt;The Looming Threat of Quantum Computing&lt;/h2&gt;
&lt;p&gt;We&apos;ve referenced Shor&apos;s algorithm many times in this series already. It can efficiently factor large numbers. This poses a significant threat to RSA, potentially rendering it insecure. The development of large-scale, fault-tolerant quantum computers is a matter of &quot;when&quot;, not &quot;if.&quot; This necessitates the urgent exploration and adoption of post-quantum cryptographic algorithms.&lt;/p&gt;
&lt;p&gt;The NIST Post-Quantum Cryptography standardization process has identified several quantum-resistant algorithms. Transitioning to these algorithms is a crucial task. It requiring careful planning, testing, and implementation. This is not a short-term project; it&apos;s a generational shift in how we secure our digital world.&lt;/p&gt;
&lt;h2 id=&quot;the-bottom-line-rsa---still-relevant-but-prepare-for-the-future&quot;&gt;The Bottom Line: RSA - Still Relevant, But Prepare for the Future!&lt;/h2&gt;
&lt;p&gt;RSA stands as a testament to the power of mathematical elegance in cryptography. We&apos;ve journeyed from its prime number foundations through practical implementations, exploring how this revolutionary algorithm transformed digital security. We however cannot ignore how quantum computing poses an existential threat via Shor&apos;s algorithm. RSA&apos;s design principles - particularly the creative use of computational asymmetry, till date continue to inspire modern cryptographic solutions. This includes many post-quantum candidates.&lt;/p&gt;
&lt;p&gt;As we look ahead to our next deep dive into Elliptic Curve Cryptography (ECC), it&apos;s worth reflecting on RSA&apos;s enduring legacy. Its elegant mathematical foundation didn&apos;t just solve the public-key cryptography challenge, but it showed us how creative thinking about computational complexity could revolutionize security. Do share your experiences with RSA and possibly it&apos;s implementations! Have you encountered interesting challenges with key sizes or padding schemes?&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: Hiding in Plain Sight with Diffie-Hellman]]></title><description><![CDATA[After exploring the elegant world of SHA-3's sponge construction in Part 5, we're shifting gears to tackle one of cryptography's most…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-hiding-in-plainsight-with-diffie-hellman</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-hiding-in-plainsight-with-diffie-hellman</guid><pubDate>Sun, 20 Oct 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;After exploring the elegant world of SHA-3&apos;s sponge construction in Part 5, we&apos;re shifting gears to tackle one of cryptography&apos;s most revolutionary breakthroughs - the Diffie-Hellman key exchange. Unlike the patent-free approach we saw with Bruce Schneier&apos;s Blowfish, DH sparked intense patent disputes that highlighted its groundbreaking importance. This fundamentally innovative protocol has been the silent guardian of our digital secrets for decades. It&apos;s been powering everything from HTTPS connections to secure messaging apps.&lt;/p&gt;
&lt;p&gt;You&apos;ve likely crossed paths with DH countless times today alone - every time you&apos;ve seen that little padlock in your browser or sent an encrypted message. But today we&apos;re ditching the surface-level chatter to dive deep into the elegant mathematics and clever design that make this protocol tick. From its mathematical foundations to its quantum-resistant future, we&apos;ll explore how this cryptographic legend continues to evolve and remain relevant in today&apos;s wild west of cybersecurity.&lt;/p&gt;
&lt;h2 id=&quot;the-genesis-of-diffie-hellman&quot;&gt;The Genesis of Diffie-Hellman&lt;/h2&gt;
&lt;p&gt;At one point cryptography was a heavily guarded realm that was primarily controlled by governments and military organizations. This was before Diffie-Hellman broke into the cryptography scene. Secure communication relied on prior shared secret keys. As you can imagine this posed a significant logistical challenge: how do you securely share these keys without them falling into the wrong hands? It was like trying to deliver a top-secret message via carrier pigeon through a flock of hungry hawks.&lt;/p&gt;
&lt;p&gt;Then came &lt;a href=&quot;https://en.wikipedia.org/wiki/Whitfield_Diffie&quot;&gt;Whitfield Diffie&lt;/a&gt; and &lt;a href=&quot;https://en.wikipedia.org/wiki/Martin_Hellman&quot;&gt;Martin Hellman&lt;/a&gt;, without any doubt two cryptographic revolutionaries who dared to dream of a different world – a world where secure communication was accessible to everyone. In their groundbreaking 1976 paper, &quot;New Directions in Cryptography,&quot; (&lt;a href=&quot;https://ieeexplore.ieee.org/document/1055638&quot;&gt;IEEE, 1976&lt;/a&gt;) they introduced the concept of public-key cryptography. This changed the face of secure communication from the ground up. They also brought in &lt;a href=&quot;https://en.wikipedia.org/wiki/Ralph_Merkle&quot;&gt;Ralph Merkle&lt;/a&gt; to the project. Ralph&apos;s work on public key distribution was fundamental but often forgotten, so sometimes this protocol is also called Diffie-Hellman-Merkle key exchange. At the time during its research and publishing there was heavy interference from the NSA who believed cryptography should be under government control. Diffie and Hellman fought to publish their revolutionary ideas and because of that we now have widely accessible secure communication.&lt;/p&gt;
&lt;p&gt;Diffie-Hellman was the first viable implementation of public-key cryptography. It enabled two parties to establish a shared secret key over an insecure channel. This was all done without ever exchanging the secret itself. This allowed for the two sides to be on the other end of the world and establish a secure form of communication over an insecure line.&lt;/p&gt;
&lt;h2 id=&quot;the-math-behind-the-magic&quot;&gt;The Math Behind the Magic&lt;/h2&gt;
&lt;p&gt;Let&apos;s get to the heart and soul of Diffie-Hellman: the intricate dance of numbers that makes secure communication possible. It does indeed involve some seriously clever math but don&apos;t worry we&apos;ll get through it together.&lt;/p&gt;
&lt;p&gt;At the center of this DH-dance-floor is &lt;em&gt;a discrete logarithm problem&lt;/em&gt;. The basis of DH starts from a computational challenge - it&apos;s relatively easy to multiply two numbers, but given a number it&apos;s difficult to get to its exact roots. With that now imagine a one-way street in the world of numbers. You can easily multiply a number by itself multiple times (aka the easy street). But try to figure out how many times it was multiplied to get to this result? That&apos;s like navigating a dance floor blindfolded (aka the hard street).&lt;/p&gt;
&lt;p&gt;How about we bring in the classical dancers of DH-World - Alice, Bob, and Eve. Picture this: Alice, wants to share a secret with Bob, but Eve is eavesdropping on every word. Difficult to deal with, I know. How do you whisper secrets in a crowded room? Enter the gurus - Diffie-Hellman.&lt;/p&gt;
&lt;p&gt;Here&apos;s the choreography:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Setting the Stage:&lt;/strong&gt; Alice and Bob agree on two public numbers: a base &apos;g&apos; and a large prime &apos;p&apos;. These are the rules of the cryptographic dance, visible to everyone, including Eve.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Secret Spin:&lt;/strong&gt; Alice now chooses a secret number &apos;a&apos; (private key), and then raises &apos;g&apos; to the power of &apos;a&apos;. From there she takes the result modulo &apos;p&apos;. This gives you her public key &apos;A&apos;. Bob does the same with his secret number &apos;b&apos;, calculating his public key &apos;B&apos;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Exchange:&lt;/strong&gt; Alice and Bob now exchange their public keys – &apos;A&apos; and &apos;B&apos; – loud and clear for everyone to hear. Eve now knows &apos;g&apos;, &apos;p&apos;, &apos;A&apos;, and &apos;B&apos;, but crucially, she doesn&apos;t know &apos;a&apos; or &apos;b&apos;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Magical Merge:&lt;/strong&gt; Alice takes Bob&apos;s public key &apos;B&apos; and raises it to the power of her secret &apos;a&apos;. She then takes the result modulo &apos;p&apos;. This gives her the shared secret &apos;s&apos;. Bob does the same with Alice&apos;s public key &apos;A&apos; and his secret &apos;b&apos;. And here&apos;s the kicker: due to the magic of modular arithmetic, both of them arrive at the exact same secret &apos;s&apos;! Amazing right!?&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img src=&quot;https://en.wikipedia.org/wiki/File:DiffieHellman.png&quot; alt=&quot;DH Key Exchange Diagram&quot;&gt;
&lt;em&gt;A simple diagram illustrating the key exchange process, showing the flow of public keys and the calculation of the shared secret. Source: &lt;a href=&quot;https://en.wikipedia.org/wiki/Diffie%E2%80%93Hellman_key_exchange&quot;&gt;Wikipedia&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Let&apos;s break down the math, step by step:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Alice calculates:&lt;/strong&gt; A = g&lt;sup&gt;a&lt;/sup&gt; mod p&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bob calculates:&lt;/strong&gt; B = g&lt;sup&gt;b&lt;/sup&gt; mod p&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alice receives B and calculates:&lt;/strong&gt; s = B&lt;sup&gt;a&lt;/sup&gt; mod p = (g&lt;sup&gt;b&lt;/sup&gt; mod p)&lt;sup&gt;a&lt;/sup&gt; mod p = g&lt;sup&gt;ab&lt;/sup&gt; mod p&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bob receives A and calculates:&lt;/strong&gt; s = A&lt;sup&gt;b&lt;/sup&gt; mod p = (g&lt;sup&gt;a&lt;/sup&gt; mod p)&lt;sup&gt;b&lt;/sup&gt; mod p = g&lt;sup&gt;ab&lt;/sup&gt; mod p&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Spot the magic, Pen-and-Teller? Both Alice and Bob arrive at g&lt;sup&gt;ab&lt;/sup&gt; mod p, the shared secret, without ever directly exchanging &apos;a&apos; or &apos;b&apos;. Eve, despite knowing &apos;g&apos;, &apos;p&apos;, &apos;A&apos;, and &apos;B&apos;, is left scratching her head, unable to efficiently calculate &apos;s&apos; without knowing either &apos;a&apos; or &apos;b&apos;. This is the power of the discrete logarithm problem in action.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;# Simplified DH example (don&apos;t use this in production, seriously!)
import random

def diffie_hellman(p, g):
    a = random.randint(2, p - 1)    # Alice&apos;s private key – shhh!
    A = pow(g, a, p)                # Alice&apos;s public key – shout it from the rooftops!

    b = random.randint(2, p - 1)    # Bob&apos;s private key – top secret!
    B = pow(g, b, p)                # Bob&apos;s public key – for everyone to see!

    s_a = pow(B, a, p)              # Alice computes the shared secret
    s_b = pow(A, b, p)              # Bob computes the shared secret

    assert s_a == s_b               # They should be the same! – Magic!
    return s_a

# Example usage
p = 23                              # A small prime (in reality, use much larger primes)
g = 5                               # A generator
shared_secret = diffie_hellman(p, g)
print(f&quot;Shared secret: {shared_secret}&quot;) # Shared secret: 1
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Needless to say, like with samples all throughout this series - this code is just a simplified illustration. Real-world implementations use industrial-strength primes and handle all sorts of edge cases.&lt;/p&gt;
&lt;h2 id=&quot;a-family-of-secrets&quot;&gt;A Family of Secrets&lt;/h2&gt;
&lt;p&gt;Don&apos;t be fooled by Diffie-Hellman, it isn&apos;t just one move. It&apos;s a whole damn dance crew where each member brings in their own unique style to the cryptographic floor. Over time we&apos;ve seen some seriously slick movers and shakers that have expanded on the original DH. We&apos;ve got bright minds to pick up DH and create variations to address specific security threats and boost performance. It&apos;s the World Dance Faceoff: you&apos;ve got the OG DH, the breakdancing duo ECDH, the security-conscious hip-hop crew Authenticated DH, and the flash mob that is Group DH. Let&apos;s meet the crew:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Elliptic Curve DH (ECDH):&lt;/strong&gt; Picture DH but as a pair of insane breakdancers who are pulling off moves you never thought possible. ECDH leverages the mind-bending math of elliptic curves (think finite fields and geometric wizardry) to achieve the same level of security. The unique selling point is that they do this with way smaller key sizes. This makes it the MVP for resource-constrained environments like IoT devices.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Authenticated DH:&lt;/strong&gt; The OG DH, bless its heart, is nothing short of a free spirit. It trusts everyone and doesn&apos;t inherently protect against man-in-the-middle (MITM) attacks. Authenticated DH comes in with the hip-hop swagger by adding an identity check to the routine. It checks against digital signature to make sure that you&apos;re really grooving with Bob and not some malicious Eve trying to steal your moves.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Group DH:&lt;/strong&gt; Sometimes a bigger party is what the routing needs and in such cases you need to share the secret handshake with the whole crew. Imagine planning a surprise flash mob. It&apos;s a surprise only if you find a way to coordinate everyone without spilling the beans. That&apos;s where Group DH walks in. It takes the core DH moves and adapts them for group communication, letting multiple parties establish a shared secret.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Beyond these headliners, there are many more like the Station-to-Station (STS) protocol. It combines DH with signatures for mutual authentication (like a choreographed duet with a signed contract). There&apos;s also the MTI/A0 key agreement protocol, a frequent performer in GSM networks. The DH dance crew is truly a diverse bunch where each member bringing their unique skills and style to the cryptographic stage.&lt;/p&gt;
&lt;h2 id=&quot;dh-dance-crew-face-off-comparing-the-moves&quot;&gt;DH Dance Crew Face-Off: Comparing the Moves&lt;/h2&gt;
&lt;p&gt;Okay, so we&apos;ve met the DH dance crew. Now let&apos;s see how they stack up against each other in a head-to-head comparison:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Original DH (The Founder)&lt;/th&gt;
&lt;th&gt;ECDH (The Breakdancing Duo)&lt;/th&gt;
&lt;th&gt;Authenticated DH (The Hip-Hop Crew)&lt;/th&gt;
&lt;th&gt;Group DH (The Flash Mob)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Key Size&lt;/td&gt;
&lt;td&gt;Larger&lt;/td&gt;
&lt;td&gt;Smaller&lt;/td&gt;
&lt;td&gt;Similar to Original DH&lt;/td&gt;
&lt;td&gt;Varies depending on the group size&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Slower&lt;/td&gt;
&lt;td&gt;Faster&lt;/td&gt;
&lt;td&gt;Slightly slower than Original DH&lt;/td&gt;
&lt;td&gt;Can be complex and slower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource Usage&lt;/td&gt;
&lt;td&gt;Higher&lt;/td&gt;
&lt;td&gt;Lower&lt;/td&gt;
&lt;td&gt;Slightly higher than Original DH&lt;/td&gt;
&lt;td&gt;Higher, especially with large groups&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;Good, but needs backup&lt;/td&gt;
&lt;td&gt;Higher (for equivalent key size)&lt;/td&gt;
&lt;td&gt;Significantly higher&lt;/td&gt;
&lt;td&gt;Good, but coordination is key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vulnerability&lt;/td&gt;
&lt;td&gt;Susceptible to MITM&lt;/td&gt;
&lt;td&gt;Susceptible to MITM&lt;/td&gt;
&lt;td&gt;Resistant to MITM&lt;/td&gt;
&lt;td&gt;Subgroup attacks, complexity issues&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;quantum-quandaries-the-future-of-dh&quot;&gt;Quantum Quandaries (The Future of DH)&lt;/h2&gt;
&lt;p&gt;We cannot escape the looming threat of quantum computing. As we&apos;ve seen many times in the past few articles of the series - these futuristic machines bring with them their mind-boggling power. This raw computational power poses a serious challenge to many cryptographic algorithms that including DH. Shor&apos;s algorithm which is a quantum algorithm for factoring large numbers can in theory crack the discrete logarithm problem. This means that it can render DH vulnerable.&lt;/p&gt;
&lt;p&gt;But fear not! Cryptographers are already working on &lt;em&gt;post-quantum cryptography&lt;/em&gt; – algorithms resistant to attacks from both classical and quantum computers. Promising contenders include lattice-based cryptography, code-based cryptography, and hash-based cryptography. These technologies are still evolving, but they represent the future of secure communication in a post-quantum world.&lt;/p&gt;
&lt;h2 id=&quot;keeping-secrets-safe-security-considerations&quot;&gt;Keeping Secrets Safe (Security Considerations)&lt;/h2&gt;
&lt;p&gt;No system is impenetrable, and DH has its vulnerabilities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Man-in-the-Middle (MITM) Attacks:&lt;/strong&gt; A sneaky attacker can intercept the public key exchange and then establish separate shared secrets with each party. This allows this attacker to relay messages in plaintext effectively eavesdropping on the conversation. Authenticated DH is your shield against this threat.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Small Subgroup Attacks:&lt;/strong&gt; If parameters aren&apos;t carefully chosen, attackers might exploit weaknesses in smaller subgroups to compromise the key exchange. This is why the seniors keep on repeating - choose your parameters wisely!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Safe Prime Selection:&lt;/strong&gt; The choice of the prime number &apos;p&apos; is crucial. Using a &quot;safe prime&quot; (a prime number where (p-1)/2 is also prime) significantly strengthens the protocol.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;dh-in-the-wild-modern-applications&quot;&gt;DH in the Wild (Modern Applications)&lt;/h2&gt;
&lt;p&gt;Diffie-Hellman is everywhere, quietly securing our digital lives:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;TLS/SSL:&lt;/strong&gt; That little padlock in your browser? DH often plays a key role in establishing a secure connection between your browser and the server.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Signal Protocol:&lt;/strong&gt; This end-to-end encrypted messaging app relies on DH for its key exchange magic.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;VPN:&lt;/strong&gt; VPNs use DH to create secure tunnels, protecting your data as it travels across the internet.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;IoT Device Pairing:&lt;/strong&gt; When you pair your smart devices to your home network, DH is often working behind the scenes to secure the connection.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;bottom-line&quot;&gt;Bottom Line&lt;/h2&gt;
&lt;p&gt;Diffie-Hellman stands as a masterpiece of cryptographic innovation by elegantly solving the key distribution problem that plagued secure communication for centuries. From its mathematical foundations in discrete logarithms to modern elliptic curve variants - we&apos;ve seen how DH&apos;s clever design enables secure key exchange over insecure channels. While quantum computing poses future challenges for all cryptographic operations, the protocol&apos;s adaptability has already spawned quantum-resistant variants. This goes a long way in ensuring its continued relevance in tomorrow&apos;s cryptographic landscape.&lt;/p&gt;
&lt;p&gt;That&apos;s not all though. Stay tuned for our next deep dive into RSA. We&apos;ll explore another revolutionary public-key cryptography system that builds upon many concepts we&apos;ve covered here. Till then remember that every time you see that padlock in your browser, you&apos;re witnessing Diffie and Hellman&apos;s elegant mathematics at work. It&apos;s working silently in protecting your digital communications. What are your thoughts on DH&apos;s evolution? Do share your experiences implementing or working with this cryptographic cornerstone!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: SHA-3 – Absorbing the Complexity, Squeezing Out Security]]></title><description><![CDATA[Fresh from our exploration of HMAC's authentication prowess in Part 4, we're diving into SHA-3. For a change, It's the quantum-resistant…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-sha3-sponge-security</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-sha3-sponge-security</guid><pubDate>Tue, 08 Oct 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Fresh from our exploration of HMAC&apos;s authentication prowess in Part 4, we&apos;re diving into SHA-3. For a change, It&apos;s the quantum-resistant powerhouse of modern cryptographic hashing. While its predecessors SHA-1 and SHA-2 laid the groundwork, SHA-3 represents a revolutionary approach with its unique sponge construction. It&apos;s a complete departure from traditional Merkle-Damgård designs.&lt;/p&gt;
&lt;p&gt;In this fifth installment of our series, we&apos;ll dissect SHA-3&apos;s innovative architecture, from its permutation-based core to its robust security properties. Whether you&apos;re building blockchain applications, implementing digital signatures, or preparing systems for the quantum computing era, understanding SHA-3&apos;s capabilities is crucial. Get ready to absorb the complexity and squeeze out the security benefits as we unravel the algorithm that&apos;s reshaping the future of cryptographic hashing!&lt;/p&gt;
&lt;h2 id=&quot;the-dawn-of-sha-3-a-preemptive-strike-against-cryptographic-monoculture&quot;&gt;The Dawn of SHA-3: A Preemptive Strike Against Cryptographic Monoculture&lt;/h2&gt;
&lt;p&gt;It is somewhat surprising how SHA-3 started off. Most other algorithms in this class started to fix something with other algorithm. Well, not this one. SHA-3&apos;s origin story begins not with a crisis, but for a change with foresight. While SHA-2 remained a stalwart in the cryptographic world, the community wisely recognized the underlying risk of relying heavily on a single family of hash functions. A classic case of having all the eggs in one basket. Imagine a digital ecosystem where one vulnerability could be enough to bring down the entire security infrastructure. This very concern prompted &lt;a href=&quot;https://www.nist.gov&quot;&gt;NIST&lt;/a&gt; to launch the &lt;a href=&quot;https://en.wikipedia.org/wiki/NIST_hash_function_competition&quot;&gt;hash function competition&lt;/a&gt; in 2007, a preemptive strike against cryptographic monoculture.&lt;/p&gt;
&lt;p&gt;The goal wasn&apos;t to replace SHA-2 but to instead promote a diverse and resilient cryptographic toolkit. From this competition, we got Keccak - the elegant sponge construction. Think of it like a sponge absorbing data: it soaks up the input and then squeezes out a unique cryptographic fingerprint (the hash). Unlike SHA-2&apos;s &lt;a href=&quot;https://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_construction&quot;&gt;Merkle-Damgård construction&lt;/a&gt;, the sponge brings to the table adaptable output lengths and the innovative extendable-output functions (XOFs) like SHAKE128 and SHAKE256. This adaptability is what ended up standing strong in the face of emerging challenges like lightweight cryptography and the looming threat of quantum computing.&lt;/p&gt;
&lt;h2 id=&quot;the-quantum-shadow-from-theoretical-threat-to-looming-reality&quot;&gt;The Quantum Shadow: From Theoretical Threat to Looming Reality&lt;/h2&gt;
&lt;p&gt;Quantum computing, is no longer a theoretical concept. For crying out loud, you have offerings like &lt;a href=&quot;https://azure.microsoft.com/en-in/products/quantum&quot;&gt;Azure Quantum Cloud Computing&lt;/a&gt;. Quantum world now casts a long shadow over the cryptographic landscape. While large-scale, fault-tolerant quantum computers are still some time away, their potential to shatter widely used algorithms like RSA and ECC is not just a theoretical concern. Quantum algorithms like Shor&apos;s algorithm for efficient factorization and discrete logarithm computation, pose a direct threat to these foundational cryptographic primitives. Even hash functions are not entirely immune to the quantum threat. Grover&apos;s algorithm can speed up brute-force attacks, making larger hash outputs necessary for robust security.&lt;/p&gt;
&lt;p&gt;This is where SHA-3&apos;s inherent quantum resistance and flexibility come into the picture. Unlike SHA-2, whose underlying mathematical structures are vulnerable to Shor&apos;s algorithm, SHA-3&apos;s security is rooted in the sponge construction. This construction is resistant to known quantum attacks. Combining this with its support for larger output sizes provides a significantly higher level of confidence in a post-quantum world. Furthermore, SHA-3&apos;s adaptability allows for adjustments to its internal parameters. Like increasing the state size which further strengthens its defenses as quantum computing technology advances. This positions SHA-3 as a vital component in the ongoing evolution of cryptography.&lt;/p&gt;
&lt;h2 id=&quot;dissecting-the-sponge-a-deep-dive-into-sha-3s-inner-workings&quot;&gt;Dissecting the Sponge: A Deep Dive into SHA-3&apos;s Inner Workings&lt;/h2&gt;
&lt;p&gt;SHA-3 which is based on the Keccak sponge construction operates on a state of 1600 bits. These bits are visualized as a 5x5x64-bit three-dimensional array, often represented as &lt;code&gt;a[x][y][z]&lt;/code&gt;. It processes data through a series of permutation rounds, each with five steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;θ (Theta):&lt;/strong&gt; Here the focus is on diffusion. Imagine each column &lt;code&gt;C[x][z]&lt;/code&gt; as a vertical line through the state. θ calculates parity for each column and then mixes this parity into two other columns. This spreads changes across the entire state quickly.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;State (simplified):
[ ][ ][ ][ ][ ]
[ ][ ][ ][ ][ ]
[ ][ ][C][ ][ ]  &amp;#x3C;- Calculate parity for this column
[ ][ ][ ][ ][ ]
[ ][ ][ ][ ][ ]

Mix parity into other columns:
[ ][X][ ][ ][X]
[ ][X][ ][ ][X]
[ ][X][C][ ][X]
[ ][X][ ][ ][X]
[ ][X][ ][ ][X]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Where &apos;X&apos; represents the columns affected by the parity of column &apos;C&apos;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ρ (Rho):&lt;/strong&gt; Next up we rotates each 64-bit lane &lt;code&gt;a[x][y]&lt;/code&gt;. If you visualize each lane as a long horizontal sequence of bits and then rotate them around, and do this by a different amount for each lane - you would have to create asymmetry and further mixing up of the state.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Lane (simplified):
[0,1,2,3,4,5,6,7]  &amp;#x3C;- Original lane

Rotated lane (example offset of 3):
[3,4,5,6,7,0,1,2]
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;π (Pi):&lt;/strong&gt; We then rearrange the lanes within each slice. Think of it as shuffling the lanes according to a specific pattern, further scattering the bits and increasing the complexity of the state.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Simplified Lane Arrangement:
[L1][L2][L3]
[L4][L5][L6]

After π:
[L4][L2][L1]
[L5][L3][L6]  &amp;#x3C;- Lanes rearranged
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;χ (Chi):&lt;/strong&gt; This step is the only non-linear one. It operates on rows within each slice. It applies a non-linear mapping based on neighboring bits, introducing crucial non-linearity for security.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Row (simplified):
[0,1,0,1,1]

After χ (example):
[1,0,1,0,0] &amp;#x3C;- Non-linear transformation based on neighbors
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;ι (Iota):&lt;/strong&gt; Finally in this step we add a round constant to a single lane, &lt;code&gt;a[0][0]&lt;/code&gt;. This constant is different for each round, preventing repetitive patterns and ensuring each round is unique.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;These five steps are repeated for 24 rounds for SHA3-224/256 or for 12 rounds in the case of SHAKE128/256. In the end, we have created a complex and highly diffusive transformation. After absorption, the next phase which is the squeezing phase extracts the variable-length hash output. SHA-3 also offers fixed-output variants: SHA3-224, SHA3-256, SHA3-384, and SHA3-512.&lt;/p&gt;
&lt;h2 id=&quot;show-me-the-code&quot;&gt;Show Me the Code!&lt;/h2&gt;
&lt;p&gt;Fortunately, you don&apos;t have to implement this all from scratch. Here&apos;s a simple example to get SHA3-256:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import hashlib

def sha3_256_hash(data):
    sha3_hash = hashlib.sha3_256(data.encode(&apos;utf-8&apos;)).hexdigest()
    return sha3_hash

my_string = &quot;This is a test string&quot;
hashed_string = sha3_256_hash(my_string)
print(f&quot;The SHA3-256 hash of &apos;{my_string}&apos; is: {hashed_string}&quot;)

# The SHA3-256 hash of &apos;This is a test string&apos; is: 22cda3a2d2053f0a81fbf89ac531f724989f2308f0a2f29d894b07b4fec0320e
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;the-security-fortress-sha-3s-robust-defenses&quot;&gt;The Security Fortress: SHA-3&apos;s Robust Defenses&lt;/h2&gt;
&lt;p&gt;A lot of SHA-3&apos;s security is rooted in its design and extensive cryptanalysis over the years. Formal security proofs and years of scrutiny support its robustness. It exhibits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Collision Resistance:&lt;/strong&gt; Finding two different inputs producing the same hash is computationally infeasible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preimage Resistance:&lt;/strong&gt; Retrieving the original input from a given hash is practically impossible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Second Preimage Resistance:&lt;/strong&gt; Finding another input with the same hash, given an input and its hash, is extremely difficult.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resistance to Length-Extension Attacks:&lt;/strong&gt; SHA-3&apos;s sponge construction inherently avoids this SHA-2 vulnerability.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;sha-2-vs-sha-3-a-tale-of-two-hashes&quot;&gt;SHA-2 vs. SHA-3: A Tale of Two Hashes&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;SHA-2&lt;/th&gt;
&lt;th&gt;SHA-3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Construction&lt;/td&gt;
&lt;td&gt;Merkle-Damgård&lt;/td&gt;
&lt;td&gt;Keccak Sponge Construction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internal Structure&lt;/td&gt;
&lt;td&gt;Based on block ciphers&lt;/td&gt;
&lt;td&gt;Based on permutation functions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantum Resistance&lt;/td&gt;
&lt;td&gt;Vulnerable&lt;/td&gt;
&lt;td&gt;More resistant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speed&lt;/td&gt;
&lt;td&gt;Generally faster&lt;/td&gt;
&lt;td&gt;Generally slower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flexibility&lt;/td&gt;
&lt;td&gt;Fixed output lengths&lt;/td&gt;
&lt;td&gt;Variable and extendable output lengths&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;performance-and-efficiency-sha-3-in-a-world-of-varying-resources&quot;&gt;Performance and Efficiency: SHA-3 in a World of Varying Resources&lt;/h2&gt;
&lt;p&gt;We unfortunately aren&apos;t getting the best of both worlds here as SHA-3 isn&apos;t known for its blazing speed. It is generally slower when compared to SHA-2. This comes from the fact that it has more complex permutation operations within its sponge construction. SHA-2 over the years has also gone through extensive optimization efforts both in terms of software and hardware. Raw speed however isn&apos;t the sole determinant of an algorithm&apos;s suitability. For many applications, the enhanced security and quantum resistance offered by SHA-3 outweigh the performance trade-off. Take for example digital signatures or blockchain implementations - the overall time impact of hashing is often negligible compared to other operations like network latency or even the disk I/O.&lt;/p&gt;
&lt;p&gt;However, in performance-critical systems processing high volumes of data, every milliseconds matter. SHA2-256 is typically 1.5-2x faster than SHA3-256. To put that in perspective - Google processes over 8.5B searches daily. Adding even a single milli-second to each search means adding over ~98.3 days or ~3.2months of computational load. All this to say that in the real world, careful benchmarking and performance testing are crucial to determining whether SHA-3&apos;s added security justifies the potential performance hit. Consider the specific requirements of your application and the available hardware resources before making any choice.&lt;/p&gt;
&lt;p&gt;This performance consideration becomes even more critical decision makers in resource-constrained environments. Think of use cases around IoT devices, mobile phones, etc. These devices often operate with limited processing power, memory, and energy. Deploying SHA-3 on such devices may not make sense as it will lead to significant performance bottlenecks and more importantly - battery drains. This is where lightweight cryptography comes into play. Look at algorithms like Ascon and SPHINCS+ which are designed specifically for resource-constrained devices. They offer a trade-off between security and efficiency. Choosing between SHA-3 and a lightweight alternative depends on the specific security requirements as well as the resources at the disposal. If security is paramount and resources are plentiful, SHA-3 is a no-brainer. In other cases, a lightweight algorithm might be the more practical choice.&lt;/p&gt;
&lt;h2 id=&quot;sha-3-in-action&quot;&gt;SHA-3 in Action&lt;/h2&gt;
&lt;p&gt;SHA-3&apos;s versatility extends beyond theoretical cryptographic discussions. It has already found practical use in diverse applications:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Digital Signatures:&lt;/strong&gt; SHA-3 plays a crucial role in digital signatures, ensuring data integrity and authenticity. By encrypting the hash of a document with a private key, the sender creates a digital signature. The recipient can then verify the signature using the sender&apos;s public key. This helps to establish that the document hasn&apos;t been tampered with and that it originated from the claimed sender. SHA-3&apos;s collision resistance is the basis of the confidence here. Forging a valid signature would require finding another document with the same hash which as we discussed is a computationally infeasible task.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blockchain and Cryptocurrency:&lt;/strong&gt; In the world of blockchain and cryptocurrencies SHA-3 plays a significant role. It secures transactions and maintains the integrity of the distributed ledger. Cryptocurrencies like Ethereum utilize SHA-3 for various cryptographic operations. This includes transaction verification as well as mining. SHA-3&apos;s preimage resistance is vital here as it prevents attackers from reversing transactions or creating fraudulent ones.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Random Number Generation (RNG):&lt;/strong&gt; SHA-3 can be used to build cryptographically secure random number generators (CSPRNGs). This is done by hashing a seed value and then repeatedly hashing the output. This gives us a sequence of unpredictable random numbers. This is essential for applications requiring high-quality randomness. Think of use cases like generating cryptographic keys, session IDs, and nonces.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Password Hashing:&lt;/strong&gt; While traditionally bcrypt and scrypt are preferred for password hashing, the shift is coming in. SHA-3 offers a robust alternative. Generally, in such cases, we make use of a sufficiently large output size and incorporate a salt (a random value unique to each password). SHA-3 can effectively protect passwords against brute-force and rainbow table attacks.&lt;/p&gt;
&lt;h2 id=&quot;the-bottom-line-sha-3s-legacy-and-future&quot;&gt;The Bottom Line: SHA-3&apos;s Legacy and Future&lt;/h2&gt;
&lt;p&gt;SHA-3 stands as a testament to cryptographic innovation. It introduces the revolutionary sponge construction that fundamentally changed how we approach hash functions. Its quantum resistance and flexible output capabilities make it particularly valuable as we face evolving security challenges. While it may not always be the fastest option, its mathematical elegance and robust security properties have earned it a crucial place in modern cryptography.&lt;/p&gt;
&lt;p&gt;The true power of SHA-3 lies not just in its technical capabilities but in how it demonstrates the importance of cryptographic diversity. We&apos;ve seen how different use cases can all make use of it. Be it anything from blockchain implementations to resource-constrained IoT devices. They all demand different trade-offs between security, performance, and flexibility and SHA-3 stands strong. SHA-3&apos;s adaptable design principles continue to influence the development of new cryptographic primitives, including many post-quantum candidates.&lt;/p&gt;
&lt;p&gt;Do share your experiences with SHA-3! Have you encountered interesting performance trade-offs? How are you balancing security requirements with resource constraints?&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: Hashing out HMAC]]></title><description><![CDATA[Continuing our journey through the cryptographic landscape in the "Insights in Plaintext" series, we're shifting gears from symmetric…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-hashing-out-hmac</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-hashing-out-hmac</guid><pubDate>Fri, 27 Sep 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Continuing our journey through the cryptographic landscape in the &quot;Insights in Plaintext&quot; series, we&apos;re shifting gears from symmetric encryption. We&apos;ll now look into the critical world of message authentication. Part 4 brings us face-to-face with HMAC. It&apos;s a cryptographic heavyweight that&apos;s been silently securing our digital world for decades now.&lt;/p&gt;
&lt;p&gt;While HMAC might not be the newest kid on the cryptographic block, its battle-tested reliability makes it the backbone of countless security protocols today. It plays vital role in securing API calls to protecting blockchain transactions. HMAC&apos;s elegant design continues to shine in our increasingly complex digital landscape. So grab your favorite debugging mug as we&apos;re about to dissect the inner workings of Hash-based Message Authentication Codes and understand why this seemingly simple construction remains a cornerstone of modern security architecture.&lt;/p&gt;
&lt;h2 id=&quot;the-hmac-origin-story&quot;&gt;The HMAC Origin Story&lt;/h2&gt;
&lt;p&gt;In the digital Wild West of the mid-90s, one of the top challenges in front of engineering teams was to secure online communications. Existing Message Authentication Codes (MACs) were either computationally expensive, relying on slower block ciphers, or suffered from security vulnerabilities, either ones that were already exploitable or could be in the next few years. This cryptographic conundrum called for a hero, and then in 1996 we witnessed HMAC riding into the town with guns blazing (metaphorically, of course). Born from the minds of cryptographic pioneers &lt;a href=&quot;https://en.wikipedia.org/wiki/Mihir_Bellare&quot;&gt;Mihir Bellare&lt;/a&gt;, &lt;a href=&quot;https://en.wikipedia.org/wiki/Ran_Canetti&quot;&gt;Ran Canetti&lt;/a&gt;, and &lt;a href=&quot;https://en.wikipedia.org/wiki/Hugo_Krawczyk&quot;&gt;Hugo Krawczyk&lt;/a&gt;, HMAC came in with a simple yet elegant solution: provide robust security all while leveraging the speed and ubiquity of hash functions. Their groundbreaking paper, &quot;&lt;a href=&quot;https://link.springer.com/chapter/10.1007/3-540-68697-5_1&quot;&gt;Keying Hash Functions for Message Authentication&lt;/a&gt;,&quot; laid the groundwork for a MAC function that was three - fast, secure, and easy to implement.&lt;/p&gt;
&lt;p&gt;The nested hashing construction is at the core of HMACs, incorporating inner and outer padding to thwart length extension attacks and other vulnerabilities that plagued earlier MAC designs. This clever approach, formalized in &lt;a href=&quot;https://datatracker.ietf.org/doc/html/rfc2104&quot;&gt;RFC 2104&lt;/a&gt; in 1997, quickly propelled HMAC to the forefront of web security. It became everyone&apos;s go-to solution for protecting everything under the sun - from financial transactions to sensitive data. HMAC’s enduring popularity after all these years is a testament to its ingenious design and the foresight of its creators. They were able to recognize the need for a practical and robust authentication mechanism in the burgeoning digital age.&lt;/p&gt;
&lt;h2 id=&quot;hmac-under-the-microscope&quot;&gt;HMAC Under the Microscope&lt;/h2&gt;
&lt;p&gt;Enough with the hand-waving. Let&apos;s get down to the brass tacks of how HMAC &lt;em&gt;actually&lt;/em&gt; operates. As we discussed earlier, it&apos;s all about nested hashing modules. But there&apos;s more to it than just slapping some padding on and calling it a day. The devil, as they say, is in the details ...or in this case - the hashing modules.&lt;/p&gt;
&lt;p&gt;Let&apos;s start our discussion with the Keys. HMAC expects a key of a certain length. If your key is shorter than the block size of the underlying hash function (e.g., 64 bytes for SHA-256), then the key is padded with zeros up to that block size. If it&apos;s longer, it&apos;s first hashed down to a smaller size such that it fits in the block size. From there if it&apos;s smaller than the block size then it&apos;s padded up to the size. All of this is to make sure that we are dealing with a consistent key size, irrespective of what we started with.&lt;/p&gt;
&lt;p&gt;Next comes the padding. We have two options at our disposal here - inner padding (&lt;code&gt;ipad&lt;/code&gt;) and outer padding (&lt;code&gt;opad&lt;/code&gt;). Both are derived from the key using a bitwise XOR operation with fixed constants: 0x36 for &lt;code&gt;ipad&lt;/code&gt; and 0x5c for &lt;code&gt;opad&lt;/code&gt;. Why these specific constants? Well, they are carefully chosen so as to have good bitwise difference properties. This helps in ensuring the security of the construction even if the underlying hash function has some weaknesses. Think of it as adding a layer of obfuscation. All of this is for one goal - making it harder for attackers to reverse-engineer the process.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import hmac
import hashlib

def hmac_calculate(key, message, blocksize=64):
    &quot;&quot;&quot;Calculate HMAC using SHA-256.&quot;&quot;&quot;

    # Ensure key and message are bytes
    if isinstance(key, str):
        key = key.encode(&apos;utf-8&apos;)
    if isinstance(message, str):
        message = message.encode(&apos;utf-8&apos;)

    # Hash the key if it exceeds the block size
    if len(key) &gt; blocksize:
        key = hashlib.sha256(key).digest()

    # Pad the key if it&apos;s shorter than the block size
    if len(key) &amp;#x3C; blocksize:
        key = key + b&apos;\x00&apos; * (blocksize - len(key))

    # Create the inner (ipad) and outer (opad) padding
    opad = bytes(x ^ 0x5c for x in key)
    ipad = bytes(x ^ 0x36 for x in key)

    # Perform inner hash
    inner_hash = hashlib.sha256(ipad + message).digest()
    # Perform outer hash
    outer_hash = hashlib.sha256(opad + inner_hash).digest()

    return outer_hash.hex()

# Example usage
key = &quot;MySuperSecretKey&quot;
message = &quot;Top Secret Message&quot;
hmac_output = hmac_calculate(key, message)
print(f&quot;HMAC: {hmac_output}&quot;)

# Verification against Python&apos;s built-in hmac library
built_in_hmac = hmac.new(
    key.encode(&apos;utf-8&apos;),
    message.encode(&apos;utf-8&apos;),
    hashlib.sha256
).hexdigest()
print(f&quot;Built-in HMAC: {built_in_hmac}&quot;)
assert hmac_output == built_in_hmac, &quot;HMAC implementation doesn&apos;t match built-in function&quot;

# $&gt; Python my-hmac-py
#       HMAC: a8da02b39f6144341be7b70adda46893255c6de31cadc44b90f6c9d02fb9bbac
#       Built-in HMAC: a8da02b39f6144341be7b70adda46893255c6de31cadc44b90f6c9d02fb9bbac
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now, for the start of the show: the nested hashing. First, the key is XORed with the &lt;code&gt;ipad&lt;/code&gt;, and then the message is concatenated to the result. This whole shebang is then hashed using the chosen hash function (SHA-256 in our example). The output of this first hash becomes the input of the second hash. For the second round, the key is XORed with the &lt;code&gt;opad&lt;/code&gt;, and then the output of the first hash is concatenated to it. Finally, this is hashed again, producing the final HMAC output.&lt;/p&gt;
&lt;p&gt;This two-step hashing process is what makes HMAC so secure. Even if an attacker finds a collision in the inner hash, they&apos;re still faced with the outer hash, which effectively acts as a one-way function protecting the secret key. Sounds simple right? It is. A simple-er yet elegant solution.&lt;/p&gt;
&lt;h2 id=&quot;padding-more-than-just-fluff&quot;&gt;Padding: More Than Just Fluff&lt;/h2&gt;
&lt;p&gt;As you can make out by now - the &lt;code&gt;ipad&lt;/code&gt; and &lt;code&gt;opad&lt;/code&gt; aren&apos;t just random voodoo. These constants (0x36 and 0x5c) are XORed with the key. Why XOR and why these specific values? Well, they&apos;re chosen such that the chances of collision the minimized, and to make sure that the inner and outer keys are different. This is done even if the original key is short or contains repeated bytes. This is very important for preventing attacks like length extension. Imagine if an attacker could append some random data to your legitimate message and then somehow calculate the new valid HMAC without knowing your key. A complete disaster, right? The padding is what helps prevent this by effectively rendering it computationally infeasible to extend the message without knowing the original key.&lt;/p&gt;
&lt;h2 id=&quot;hmac-security-from-street-smarts-to-mathematical-muscle&quot;&gt;HMAC Security: From Street Smarts to Mathematical Muscle&lt;/h2&gt;
&lt;p&gt;Let&apos;s now turn our attention to security. In the world of cryptography, paranoia is not just a virtue; it&apos;s a damn necessity. Over the years HMAC has proven itself to earn its reputation as a robust MAC. Let&apos;s look at both the practical security considerations and the theoretical underpinnings that make it so tough to crack.&lt;/p&gt;
&lt;p&gt;From a practical standpoint, HMAC boasts some impressive defenses:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Length Extension Attacks:&lt;/strong&gt; Thanks to the clever use of inner and outer padding, HMAC is immune to length extension attacks. This style of attack has plauged many of the more simpler hash-based MACs. Those simple paddings effectively prevent an attacker from appending data to a message and calculating a valid MAC without knowing the secret key.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Timing Attacks:&lt;/strong&gt; Timing attacks exploit the fact that cryptographic operations can take slightly different amounts of time depending on the data being processed. At times it&apos;s mere CPU clock cycles that are enough to give away the secrets. To mitigate this, constant-time comparison functions are essential when verifying HMACs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Collision Resistance:&lt;/strong&gt; HMAC inherits the collision resistance of the underlying hash function. So, if you&apos;re using a strong hash function like SHA-256 or, even SHA-3, you&apos;re in good shape. The topic is further addressed with the nested hashing construction, making finding collisions even harder.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key Size:&lt;/strong&gt; The bigger, the better. A general guideline is to use a key at least as long as the output of the hash function. A longer key increases the computational effort required to brute-force it, making your HMAC more resistant to attacks.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;HMAC just keeps on giving. For the true crypto nerds, HMAC&apos;s security isn&apos;t just based on empirical testing and real-world observations. We also have the mighty mathematicians backing it! As long as we can make sure that the underlying hash function behaves like a pseudorandom function (PRF), formal security proofs show that forging an HMAC is computationally infeasible. A PRF outputs values that are indistinguishable from random to an attacker, even if they know the input (except, of course, for the secret key). This PRF assumption allows us to mathematically prove the security of HMAC, giving us a high level of confidence in its overall robustness.&lt;/p&gt;
&lt;h2 id=&quot;hmac-in-the-real-world-practical-considerations-and-applications&quot;&gt;HMAC in the Real World: Practical Considerations and Applications&lt;/h2&gt;
&lt;p&gt;Okay, we&apos;ve covered the theoretical muscle and security properties of HMAC. But how does this all play out in practice? That&apos;s where it should shine right? Well, let&apos;s explore some practical considerations for implementation and deployment, along with a look at HMAC&apos;s wide range of applications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Choosing the Right Ingredients:&lt;/strong&gt; Selecting the appropriate hash function is crucial for HMAC&apos;s security and performance. SHA-256 and SHA-3 are generally solid choices. Goes without saying but the key management is equally critical. Store your keys securely, using robust key management systems, and always rotate them regularly. And please, for the love of all that is holy, don&apos;t hardcode your keys directly into your code!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Derivation:&lt;/strong&gt; Where do those secret keys come from? Are they really secret? Key Derivation Functions (KDFs) are the answer. KDFs take a high-entropy input (like a password or a random seed) and churn out a strong, uniformly distributed key suitable for cryptographic use. Popular KDFs include &lt;a href=&quot;https://en.wikipedia.org/wiki/PBKDF2#:~:text=PBKDF2%20applies%20a%20pseudorandom%20function,cryptographic%20key%20in%20subsequent%20operations.&quot;&gt;PBKDF2&lt;/a&gt;, &lt;a href=&quot;https://en.wikipedia.org/wiki/Bcrypt&quot;&gt;bcrypt&lt;/a&gt;, &lt;a href=&quot;https://en.wikipedia.org/wiki/Scrypt&quot;&gt;scrypt&lt;/a&gt;, and &lt;a href=&quot;https://en.wikipedia.org/wiki/Argon2&quot;&gt;Argon2&lt;/a&gt;. Remember - the best choice is the one you choose, and then the next moment it becomes the one you didn&apos;t.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Beyond HMAC:&lt;/strong&gt; While HMAC is the undisputed king of the MAC hill, it&apos;s not the only name in town, far from it. Other contenders like CMAC (Cipher-based Message Authentication Code) and &lt;a href=&quot;https://en.wikipedia.org/wiki/Poly1305&quot;&gt;Poly1305&lt;/a&gt; offer alternatives, each with its own strengths and weaknesses. However, HMAC&apos;s simplicity, widespread availability, and extensive vetting make it a compelling default choice for many applications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Advanced Attacks and Staying Ahead of the Curve:&lt;/strong&gt; No cryptographic system is ever impenetrable. While HMAC is generally very secure, advanced attacks like &lt;a href=&quot;https://en.wikipedia.org/wiki/Side-channel_attack&quot;&gt;side-channel&lt;/a&gt; attacks, exploiting information leakage through timing, power consumption, or electromagnetic emissions, can pose a threat that you will find difficult to design for. Vulnerabilities in the underlying hash function can also weaken HMAC&apos;s security. Staying up-to-date on the latest security research and best practices is paramount.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HMAC in Action:&lt;/strong&gt; HMAC is ubiquitous in today&apos;s interconnected world - you may not realize it. It safeguards API authentication, ensures data integrity, secures blockchain transactions, and plays a crucial role in protocols like SSL/TLS. It&apos;s the silent guardian, the watchful protector, working tirelessly behind the scenes to keep your data safe from those pesky eavesdroppers and malicious actors.&lt;/p&gt;
&lt;h2 id=&quot;the-bottom-line-hmacs-enduring-legacy-in-modern-security&quot;&gt;The Bottom Line: HMAC&apos;s Enduring Legacy in Modern Security&lt;/h2&gt;
&lt;p&gt;Our deep dive into HMAC has revealed why this elegant cryptographic construction continues to dominate the authentication landscape. Its clever combination of nested hashing and key management provides robust security. It does so while remaining remarkably simple to implement. From its mathematical foundations to practical Python implementations, we saw how HMAC effectively addresses critical security challenges. This becomes crucial in domains like API authentication, blockchain transactions, and secure communications.&lt;/p&gt;
&lt;p&gt;Stay tuned for our next article, where we&apos;ll explore SHA-3 and its revolutionary sponge construction. We&apos;ll also see why it&apos;s a perfect complement to HMAC in building comprehensive cryptographic solutions. Until then, remember that HMAC&apos;s enduring success stems from its elegant simplicity: sometimes the most effective solutions are built with simple fundamentals rather than chasing complexity for complexity&apos;s sake.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[JWT Deep Dive: Beyond the Basics - Unmasking Claim Types and Their Secrets]]></title><description><![CDATA[Decoding the JWT Enigma We've looked a fair bit at JWT's in the last few posts. I want to close this round of discussion by talking about…]]></description><link>https://mayankraj.com/blog/jwt-claims-types</link><guid isPermaLink="false">https://mayankraj.com/blog/jwt-claims-types</guid><pubDate>Wed, 26 Jun 2024 18:30:00 GMT</pubDate><content:encoded>&lt;h2 id=&quot;decoding-the-jwt-enigma&quot;&gt;Decoding the JWT Enigma&lt;/h2&gt;
&lt;p&gt;We&apos;ve looked a fair bit at JWT&apos;s in the last few posts. I want to close this round of discussion by talking about JWTs claims. if you have used JWTs, you&apos;ve danced with them in authentication flows. The real question then is - are you truly unlocking their full potential? Beyond the familiar &lt;code&gt;sub&lt;/code&gt; and &lt;code&gt;exp&lt;/code&gt; claims, not so far away lies a rich world of customization just waiting to be explored. We&apos;re going to crack the code on JWT claims, revealing their power for fine-grained authorization, efficient data transmission, and much more.&lt;/p&gt;
&lt;h2 id=&quot;the-usual-suspects-standard-claims&quot;&gt;The Usual Suspects: Standard Claims&lt;/h2&gt;
&lt;p&gt;Before we even attempt to dive into the deep end, let&apos;s revisit the standard claims you&apos;ll encounter in almost every JWT:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;iss&lt;/code&gt; (Issuer):&lt;/strong&gt; Trace to the token&apos;s origin. A critical block in multi-tenant systems and when integrating with external authentication providers. Think of it as the token&apos;s return home address.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;sub&lt;/code&gt; (Subject):&lt;/strong&gt; The star of the show – the user or entity the token was issued for. This is usually a unique identifier.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;aud&lt;/code&gt; (Audience):&lt;/strong&gt; The intended recipient. This prevents token misuse by unintended services by scoping it appropriately.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;exp&lt;/code&gt; (Expiration Time):&lt;/strong&gt; The token&apos;s expiry date. Everything has a limited time, and the expiry claim makes sure that JWT do so as well. They prevent tokens from living forever and becoming a security liability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;nbf&lt;/code&gt; (Not Before):&lt;/strong&gt; The token&apos;s activation time. Useful for scheduling access or delayed functionality. Think of it as a post-dated check.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;iat&lt;/code&gt; (Issued At):&lt;/strong&gt; The token&apos;s birth timestamp. Helpful for tracking its age and detecting potential issues. Like a fine wine, some tokens get better with age (just kidding!).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code&gt;jti&lt;/code&gt; (JWT ID):&lt;/strong&gt; A unique identifier for the specific token. Helpful in preventing replay attacks, ensuring each token is a one-time-use ticket.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are the foundation, but the real magic happens when the capability of JWT claims are used to drive application use cases.&lt;/p&gt;
&lt;h2 id=&quot;claim-customization-unleashing-the-power-of-jwts&quot;&gt;Claim Customization: Unleashing the Power of JWTs&lt;/h2&gt;
&lt;p&gt;Here&apos;s where things start to get exciting. JWTs allow you to define custom claims allowing you to add rich contextual information directly into the token itself. This unlocks powerful capabilities for, but not limited to use cases around authorization, personalization, and streamlined data access.&lt;/p&gt;
&lt;h3 id=&quot;roles-and-permissions-fine-grained-access-control&quot;&gt;Roles and Permissions: Fine-Grained Access Control&lt;/h3&gt;
&lt;p&gt;Imagine a user with multiple roles, like &quot;admin&quot; and &quot;editor.&quot; You can embed these roles directly into the JWT:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;sub&quot;: &quot;12345&quot;,
  &quot;roles&quot;: [&quot;admin&quot;, &quot;editor&quot;],
  &quot;permissions&quot;: [&quot;create_article&quot;, &quot;edit_article&quot;, &quot;publish_article&quot;]
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This reduces the database round-trips for authorization as now your application can directly check these claims and grant or deny access based on the user&apos;s roles and permissions. This not only improves performance but also reduces latency, all this while also reducing the load on the overall system.&lt;/p&gt;
&lt;h3 id=&quot;user-context-personalization-and-beyond&quot;&gt;User Context: Personalization and Beyond&lt;/h3&gt;
&lt;p&gt;In a multi-tenant application, include the &lt;code&gt;tenant_id&lt;/code&gt; directly in the JWT:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;sub&quot;: &quot;12345&quot;,
  &quot;tenant_id&quot;: &quot;acme_corp&quot;,
  &quot;preferences&quot;: {
    &quot;theme&quot;: &quot;dark&quot;,
    &quot;language&quot;: &quot;en&quot;
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This not only simplifies tenant-specific logic but also allows for fine grained user experience personalization. Imagine pre-loading user preferences, UI settings, or even A/B testing flags - all within the token. Think of the last time you saw a web page load in light mode only for it to automatically switch to dark mode because you had set a preference for the dark mode. Well now your users will never have to experience it.&lt;/p&gt;
&lt;h3 id=&quot;data-minimization-privacy-by-design&quot;&gt;Data Minimization: Privacy by Design&lt;/h3&gt;
&lt;p&gt;Never ever overshare! Instead of sending the entire user profile with every request, include only the essential data relevant to the current scenario in the JWT:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &quot;sub&quot;: &quot;12345&quot;,
  &quot;username&quot;: &quot;johndoe&quot;,
  &quot;email&quot;: &quot;john.doe@example.com&quot;,
  &quot;profile_picture&quot;: &quot;https://example.com/avatars/johndoe.png&quot;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This minimizes the amount of sensitive data in transit which in turn helps in reducing the impact of potential breaches. To make the deal even more sweet, it&apos;s more efficient than sending large payloads with every request. This requires careful planning ahead of time to make sure everything needed for the job is availble in the JWT, and nothing more.&lt;/p&gt;
&lt;h2 id=&quot;walking-the-tightrope-security-considerations-and-best-practices&quot;&gt;Walking the Tightrope: Security Considerations and Best Practices&lt;/h2&gt;
&lt;p&gt;Custom claims are powerful, but like everything powerful they come with responsibilities. Remember - never treat your JWTs like a digital junk drawer. Keep them lean, mean, and secure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Size Matters:&lt;/strong&gt; Large JWTs impact overall performance. Include only the essential data. For everything else - consider using a dedicated caching mechanism.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sensitive Data Alert:&lt;/strong&gt; Needless to say - avoid putting highly sensitive data (e.g., credit card numbers, social security numbers) directly into the JWT unless you have robust encryption in place. Think end-to-end encryption or payload encryption, as JWT claims are by themselves visible (but not verifiable).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Validation is Key:&lt;/strong&gt; Always validate &lt;em&gt;every&lt;/em&gt; claim on the server-side. Don&apos;t ever trust the client-provided data. Validate data types, formats, and ranges to prevent man-in-the-middle styled attacks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Expiration Dates Are Your Friends:&lt;/strong&gt; Use appropriate, short-lived expiration times to limit the damage for the inevitable event of token being compromised. Implement refresh tokens for long-lived sessions.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;signing-ceremony-choosing-the-right-algorithm&quot;&gt;Signing Ceremony: Choosing the Right Algorithm&lt;/h2&gt;
&lt;p&gt;The security of the JWT token relies heavily on signing algorithms, or I may go as far as to say - only on the signing algorithm. These algorithms ensure the integrity as well as the authenticity of the token. Here&apos;s a breakdown of the most common choices:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;HS256 (HMAC with SHA-256):&lt;/strong&gt; Fast and efficient, uses a shared secret key. Great for internal services where key management is relatively straightforward. However, if the secret is compromised, all tokens become vulnerable. Think of one master key for all locks in the nuclear facility.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RS256 (RSA with SHA-256):&lt;/strong&gt; More secure than HS256, uses a public/private key pair. The private key signs the token, and the public key verifies it. Suitable for scenarios where you have the use case to share the verification key publicly, such as with external APIs or third-party services. It&apos;s slower than HS256 but offers better security and the option to share responsibility.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ES256 (ECDSA with SHA-256):&lt;/strong&gt; Provides strong security with smaller key sizes compared to RSA - a real win-win. A great alternative for resource-constrained environments. Elliptic curve cryptography makes it performant while maintaining a high level of security.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Choosing the right algorithm depends on your specific needs - security requirements, performance needs, along with key management capabilities. For highly sensitive data or external APIs, RS256 or ES256 are generally preferred. For internal services, HS256 could be a viable option given that you keep the secret ..well ...a secret. In more sensitive use cases you could use a hardware security module (HSM) for enhanced key security.&lt;/p&gt;
&lt;h2 id=&quot;real-world-adventures-jwts-in-enterprise-applications&quot;&gt;Real-World Adventures: JWTs in Enterprise Applications&lt;/h2&gt;
&lt;p&gt;Let&apos;s also look at some real-world scenarios where JWTs take the spotlight:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;E-commerce:&lt;/strong&gt; JWTs often carry product entitlements, seamlessly controlling access to premium features as well as content. Imagine unlocking exclusive content based on a user&apos;s subscription level, all from the data encoded within the JWT itself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Healthcare:&lt;/strong&gt; Securely transmit patient identifiers and consent information within JWTs, while maintaining integrity. Thus streamlining data access for authorized personnel while adhering to HIPAA regulations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Single Sign-On (SSO):&lt;/strong&gt; Enable smooth access to multiple applications within an organization thus allowing users to log in once and access various services without re-authenticating.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Microservices Communication:&lt;/strong&gt; Secure communication between microservices using JWTs for service-to-service authentication and authorization. This enhances security and decoupling within a microservices architecture.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;jwt-toolkit-libraries-and-tools&quot;&gt;JWT Toolkit: Libraries and Tools&lt;/h2&gt;
&lt;p&gt;JWT is not experimental by any stretch of imagination. You don&apos;t have to figure the details out by yourself. Numerous libraries and tools simplify JWT management:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Auth0:&lt;/strong&gt; A comprehensive platform for authentication and authorization, with robust JWT support and extensive documentation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Firebase Auth:&lt;/strong&gt; Another great option with built-in JWT capabilities, especially for mobile and web applications. It integrates seamlessly with other Firebase services.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Language-Specific Libraries:&lt;/strong&gt; From Python&apos;s &lt;code&gt;PyJWT&lt;/code&gt; to Node.js&apos;s &lt;code&gt;jsonwebtoken&lt;/code&gt; and Java&apos;s &lt;code&gt;jjwt&lt;/code&gt;, you&apos;ll find a library for virtually every language. These libraries provide convenient functions for creating, signing, and verifying JWTs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-bottom-line-jwts-are-your-secret-weapon&quot;&gt;The Bottom Line: JWTs Are Your Secret Weapon&lt;/h2&gt;
&lt;p&gt;JWTs while often associated with just authentication tokens, they are more than that. They&apos;re versatile tools for secure data transmission, fine-grained authorization, and efficient data management. By understanding claims, signing algorithms, and best practices, you can build more efficient, secure, and personalized applications.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: A Deep Dive in the Sea with Blowfish]]></title><description><![CDATA[Welcome to Part 3 of our "Insights in Plaintext" series! After exploring the rise and fall of DES/3DES in our previous deep dive, we're…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-deep-dive-with-blowfish</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-deep-dive-with-blowfish</guid><pubDate>Tue, 18 Jun 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Welcome to Part 3 of our &quot;Insights in Plaintext&quot; series! After exploring the rise and fall of DES/3DES in our previous deep dive, we&apos;re turning our attention to Blowfish. It&apos;s a cipher that revolutionized cryptography by championing open-source principles. Think of it as that trusty old university hoodie you still keep around - while you might not wear it daily anymore, but it represents a pivotal moment in your journey. Just as DES taught us about the dangers of limited key sizes, Blowfish has its own crucial lessons about open security and algorithmic evolution. So grab your coffee, put on your nostalgia goggles, and let&apos;s dive into this game-changing algorithm that bridged the gap between legacy ciphers and modern encryption.&lt;/p&gt;
&lt;h2 id=&quot;schneiers-brainchild-open-source-before-it-was-cool&quot;&gt;Schneier&apos;s Brainchild: Open Source Before It Was Cool&lt;/h2&gt;
&lt;p&gt;Ready!? Picture this: it&apos;s the early 90s. Bell bottoms are a thing, the internet is still a baby and cryptography? Well, let&apos;s just say it was a different world. DES, the reigning champ, was starting to show cracks, and proprietary algorithms, shrouded in secrecy, were the norm. Imagine a world where you couldn&apos;t peek under the hood. This is when security relied on obscurity rather than robust design. It was a bit like buying a car and not being allowed to lift the hood to see the engine. Sketchy, right?&lt;/p&gt;
&lt;p&gt;Into this scene strides &lt;a href=&quot;https://www.schneier.com&quot;&gt;Bruce Schneier&lt;/a&gt;, a security guru with a rebellious streak. He believed in open cryptography and in the power of community scrutiny. He saw the limitations of closed-source security and unlike many decided to do something about it. Blowfish, released in 1993, was his answer – a fast, free, and unapologetically open-source encryption algorithm. It was a bold move, a direct challenge to the established order.&lt;/p&gt;
&lt;p&gt;You may ask why did Schneier create Blowfish? Several factors fueled its development:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Limitations of DES:&lt;/strong&gt; DES, with its 56-bit key, was becoming increasingly vulnerable to brute-force attacks. The US government&apos;s involvement also raised concerns about potential backdoors (never gets old, does it?).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost and Availability of Cryptography:&lt;/strong&gt; Existing alternatives were often expensive and encumbered by patents, limiting access for developers and researchers. Schneier wanted to democratize cryptography, making strong encryption available to everyone.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Philosophy of Openness:&lt;/strong&gt; Schneier firmly believed that open scrutiny leads to stronger security. By exposing the algorithm to public review, he invited the entire cryptographic community to find and fix potential weaknesses.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Open source was not the only selling point of Blowfish; it was designed for speed and simplicity. Schneier prioritized performance, especially in software implementations. This made it suitable for a wide range of applications, from securing email communications to protecting hard drives. Its compact size made it particularly attractive for resource-constrained environments, a crucial factor in the early days of computing.&lt;/p&gt;
&lt;p&gt;It would be an understatement to say that the impact of Blowfish&apos;s open-source nature was profound. It encouraged collaboration and peer review, which in turn ended up accelerating the advancement of cryptographic research. It demonstrated that security through obscurity is a fallacy, and that open algorithms, subjected to rigorous analysis, can be far more secure. Blowfish became a textbook example of how to do cryptography right. This ended up influencing the design and development of countless ciphers that followed it.&lt;/p&gt;
&lt;h2 id=&quot;under-the-hood-gears-and-gadgets&quot;&gt;Under the Hood: Gears and Gadgets&lt;/h2&gt;
&lt;p&gt;Alright, out from history, into the maths class. Let&apos;s roll up our sleeves and get into the nitty-gritty of how Blowfish actually works.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Variable Key Length (32-448 bits):&lt;/strong&gt; Yes! This flexibility allows you to fine-tune the security level to match your needs and resources. A shorter key might be acceptable for low-risk scenarios or resource-constrained devices, while a longer key provides greater protection against brute-force attacks. Remember that with longer key length comes more sophisticated key-management responsibility.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;64-bit Block Size:&lt;/strong&gt; This block size, while efficient for its time, is now a major security vulnerability due to the Birthday Paradox. It&apos;s a trade-off between performance and security that hasn&apos;t aged well.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Feistel Network&lt;/strong&gt; This is where the real magic happens. Imagine a cryptographic dance floor where data is constantly being paired, shuffled, and transformed. The Feistel network is a series of rounds, typically 16 in Blowfish, where the input data is split into two halves. One half undergoes a transformation with the key, and the result is then XORed with the other half. Then the two swap places, and the process repeats. This repeated shuffling and XOR operations end up creating a strong avalanche effect wherein relatively small changes in the input lead to significantly large and measurable changes in the output.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Scheduling&lt;/strong&gt; Why limit yourself to a small key eh? Key scheduling takes your initial key which is generally in the range of 32-448 bits, and expands it into a much larger array of subkeys. This array could be a whopping 4168 bytes! These subkeys are then used in the rounds of the Feistel network, injecting key material throughout the encryption process. The key scheduling itself involves complex operations using the digits of Pi and iterative encryption of a zero block. This seemingly random process was put in place to ensure that the subkeys are thoroughly derived from the original key, making it difficult to reverse-engineer the key from the subkeys.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;P-array and S-boxes&lt;/strong&gt; The P-array is a set of 18 32-bit subkeys used in the key scheduling process and during encryption. The S-boxes are four substitution boxes, each containing 256 32-bit entries. These S-boxes play a fundamental role in introducing non-linearity into the encryption process. They act as lookup tables, mapping 8-bit input values to 32-bit output values. The contents of both the P-array and the S-boxes are initialized using the digits of Pi and then modified using the key material during the key scheduling process. Why Pi? Maybe because it&apos;s a number that never stops giving!&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;lets-get-our-hands-dirty-code-time&quot;&gt;Let&apos;s Get Our Hands Dirty: Code Time!&lt;/h2&gt;
&lt;p&gt;Here&apos;s a slightly more detailed look at the F-function, incorporating the S-boxes. Again, this is for educational purposes only, and not to be used even in &quot;dev-envs&quot;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;def f(x, s_boxes):
    # Split the 32-bit input into four 8-bit bytes
    x1 = (x &gt;&gt; 24) &amp;#x26; 0xFF                               # Extract the bytes individually
    x2 = (x &gt;&gt; 16) &amp;#x26; 0xFF
    x3 = (x &gt;&gt; 8) &amp;#x26; 0xFF
    x4 = x &amp;#x26; 0xFF

    # The core F-function calculation
    result = (s_boxes[0][x1] + s_boxes[1][x2]) % 2**32  # Modulo for 32-bit overflow
    result ^= s_boxes[2][x3]                            # XOR operation
    result = (result + s_boxes[3][x4]) % 2**32

    return result
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This function starts with a 32-bit input &lt;code&gt;x&lt;/code&gt; and the pre-initialized S-boxes. It splits &lt;code&gt;x&lt;/code&gt; into four 8-bit bytes, which are then used as indices in the S-boxes. The resulting values from the S-boxes are combined using addition and XOR operations, with modulo operations ensuring the result stays within 32 bits. This seemingly simple function plays a crucial role in creating complex transformations within the Feistel network.&lt;/p&gt;
&lt;h2 id=&quot;the-birthday-paradox-its-not-about-cake&quot;&gt;The Birthday Paradox: It&apos;s Not About Cake&lt;/h2&gt;
&lt;p&gt;Unfortunately, the Birthday Paradox isn&apos;t about some shared cake; it&apos;s about the surprising probability of shared birthdays. It deals with the probability that in a room of just 23 people, there&apos;s more than a 50% chance that two share a birthday! That&apos;s the real surprise right!? This counterintuitive result stems from the fact that we&apos;re looking for &lt;em&gt;any&lt;/em&gt; two people to match, not one specific birthday. With more people, the number of possible pairs grows faster and so does the overall probability. This same principle applies to cryptography. Blowfish&apos;s 64-bit block size means there are 2^64 possible ciphertext blocks. While this seems like a HUGE number, the Birthday Paradox tells us that collisions become likely after encrypting roughly the square root of that, around 2^32 blocks. That&apos;s just 4GB of data or the images from your last vacation. These collisions can leak information about the plaintext, weakening the encryption.&lt;/p&gt;
&lt;h2 id=&quot;speed-demons-blowfish-vs-the-new-kids&quot;&gt;Speed Demons: Blowfish vs. The New Kids&lt;/h2&gt;
&lt;p&gt;Now to the real question - how does Blowfish stack up against the competition?&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cipher&lt;/th&gt;
&lt;th&gt;Key Setup&lt;/th&gt;
&lt;th&gt;Encryption Speed&lt;/th&gt;
&lt;th&gt;Block Size&lt;/th&gt;
&lt;th&gt;Key Size&lt;/th&gt;
&lt;th&gt;Security&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blowfish&lt;/td&gt;
&lt;td&gt;Slow&lt;/td&gt;
&lt;td&gt;Fast (software)&lt;/td&gt;
&lt;td&gt;64 bits&lt;/td&gt;
&lt;td&gt;32-448 bits&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Twofish&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;128 bits&lt;/td&gt;
&lt;td&gt;128-256 bits&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AES&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;Very Fast&lt;/td&gt;
&lt;td&gt;128 bits&lt;/td&gt;
&lt;td&gt;128, 192, 256&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Blowfish&apos;s key setup is painfully slow. AES takes the top spot for raw encryption speed, thanks to widespread hardware-level acceleration. Twofish sits comfortably in the middle. This comparison is here to paint the general idea, and as you can imagine - specialized hardware will always speed things up.&lt;/p&gt;
&lt;h2 id=&quot;beyond-the-basics-modes-of-operation&quot;&gt;Beyond the Basics: Modes of Operation&lt;/h2&gt;
&lt;p&gt;Choosing the right mode of operation is very important. Just like you wouldn&apos;t use a hammer to screw in a lightbulb, the wrong implementation can make even the best cryptographic algorithm fail. Here&apos;s a quick rundown of the options we have:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ECB (Electronic Codebook):&lt;/strong&gt; The simplest but (to no one&apos;s surprise) the least secure. Encrypting the same block always produces the same ciphertext, making patterns visible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CBC (Cipher Block Chaining):&lt;/strong&gt; Each block is XORed with the previous ciphertext block before encryption. This hides patterns but requires an Initialization Vector (IV) to get the first block rolling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CTR (Counter):&lt;/strong&gt; Turns a block cipher into a stream cipher. Fast and allows for parallel processing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GCM (Galois/Counter Mode):&lt;/strong&gt; Provides both confidentiality and authentication. A popular choice for modern applications.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Each mode has its trade-offs. Choose wisely!&lt;/p&gt;
&lt;h2 id=&quot;the-quantum-menace-is-crypto-doomed&quot;&gt;The Quantum Menace: Is Crypto Doomed?&lt;/h2&gt;
&lt;p&gt;Quantum computing is the elephant in the room, only that it&apos;s 20x the size of a normal elephant. While still in its early stages, it has the potential to break many of our current state-of-the-art cryptographic algorithms. This includes Blowfish. Shor&apos;s algorithm, a quantum algorithm, could efficiently factor large numbers and compute discrete logarithms, rendering RSA and ECC useless. Symmetric-key algorithms like Blowfish are relatively less vulnerable, but they are prone to brute force attacks. And boy are quantum computers fast. The key sizes of these algorithms would need to increase significantly to resist brute-force attacks from quantum computers. It&apos;s like preparing for a hurricane – we don&apos;t know exactly when it will hit, but we need to start building our defenses now. Post-quantum cryptography is an active area of research, exploring new algorithms that can withstand the quantum onslaught.&lt;/p&gt;
&lt;h2 id=&quot;blowfish-today-a-legacy-systems-best-friend&quot;&gt;Blowfish Today: A Legacy System&apos;s Best Friend?&lt;/h2&gt;
&lt;p&gt;Blowfish isn&apos;t the shiny new toy on the shelf anymore. AES and Twofish (Blowfish&apos;s beefier successor) have now become the go-to choices. But Blowfish still lurks in legacy systems, especially in embedded ones with limited resources - maybe one day our dishwashers will have hardware to make them smart. If you&apos;re dealing with one of these, migration is a very delicate dance. You have to balance security risks with the cost and disruption of an overhaul. It&apos;s like renovating an old house - you want to modernize it in place, but you also definitely don&apos;t want the whole thing to collapse.&lt;/p&gt;
&lt;h2 id=&quot;the-legacy-of-open-cryptography&quot;&gt;The Legacy of Open Cryptography&lt;/h2&gt;
&lt;p&gt;Blowfish stands strong as a pivotal milestone in cryptographic history. It was successful in marking the transition from closed, proprietary algorithms to open, community-reviewed security solutions. It&apos;s the bases of trusted security today. While AES and Twofish have superseded it for most of the modern applications, Blowfish still holds it&apos;s place. It&apos;s revolutionary approach to open-source security and its elegant design principles continue to influence how we develop and evaluate encryption algorithms today. Its journey from cutting-edge innovation to legacy system staple offers crucial insights into the evolution of security standards and the importance of forward-thinking design.&lt;/p&gt;
&lt;p&gt;Ready to explore more modern encryption algorithms? Join us in the next article where we&apos;ll examine HMAC and its crucial role in modern authentication systems.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: Decrypting the Legacy of DES and 3DES]]></title><description><![CDATA[Welcome to Part 2 of our "Insights in Plaintext" series! After laying the groundwork in our introduction to cryptography, we're diving deep…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-legacy-of-des-3des</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-legacy-of-des-3des</guid><pubDate>Mon, 06 May 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Welcome to Part 2 of our &quot;Insights in Plaintext&quot; series! After laying the groundwork in our introduction to cryptography, we&apos;re diving deep into our first case study: DES and 3DES. These algorithms once reigned supreme but have now been standing as fascinating monuments to the never-ending march of technical progress. Not so long ago, these ciphers were the guardians of sensitive data all over the world. Today they serve as crucial cautionary tales about complacency in cryptography. So, grab your virtual (or real) coffee mug and let&apos;s explore how these pioneering encryption standards shaped the security landscape we covered in Part 1.&lt;/p&gt;
&lt;h2 id=&quot;from-bell-labs-to-bit-rot-a-history-lesson&quot;&gt;From Bell Labs to Bit Rot: A History Lesson&lt;/h2&gt;
&lt;p&gt;Our story starts in the 1970s, a decade of groovy music, now questionable fashion choices, and the birth of modern cryptography. The powerhouse IBM, in its infinite wisdom (or perhaps with a nudge from the NSA), developed the Data Encryption Standard (DES). YES! It was the IBM. DES was a block cipher designed to protect sensitive government and commercial data. It was designed with a 56-bit key, a seemingly insurmountable barrier at the time, which gave a false sense of security. Remember this was the time when 800Kb floppy disks were state-of-the-art. Banks, eager to safeguard their customers&apos; financial transactions, embraced DES with open arms. Little did they know that the seeds of its demise were already sown.&lt;/p&gt;
&lt;p&gt;What followed was the rise of distributed computing and Moore&apos;s Law. The technological advancement gradually chipped away at DES&apos;s armor. What once seemed unbreakable began to show cracks. The final blow came in 1998, with the infamous &lt;a href=&quot;https://en.wikipedia.org/wiki/EFF_DES_cracker&quot;&gt;EFF&apos;s Deep Crack&lt;/a&gt; project, a crowdsourced effort to break DES. This sent shockwaves through the security community. One thing became painfully clear: Fort Knox was made of cardboard, and the chihuahua was asleep on the job.&lt;/p&gt;
&lt;p&gt;The scramble for a solution led to the creation of 3DES, or Triple DES. It was a band-aid attempt to bolster DES&apos;s security without completely reinventing the wheel. Why? Because we didn&apos;t have the time to do so back then (more on this later). 3DES, as the name suggests, applies the DES algorithm three times, effectively increasing the key size to 112 or 168 bits. It was always a temporary fix, never a cure. An intermediate solution to buy us more time to move to something else. While 3DES offered improved security compared to its predecessor, it also ended up inheriting DES&apos;s fundamental limitations. All of this while adding the burden of significantly increased computational overhead.&lt;/p&gt;
&lt;h2 id=&quot;under-the-hood-gears-and-gizmos&quot;&gt;Under the Hood: Gears and Gizmos&lt;/h2&gt;
&lt;p&gt;DES, at its heart, is a &lt;a href=&quot;https://en.wikipedia.org/wiki/Feistel_cipher&quot;&gt;Feistel network&lt;/a&gt;, a structure that involves breaking up the data block into two halves and then passing them through a series of intricate transformations. While it may look so this isn&apos;t just some random shuffling of bits; it&apos;s a carefully orchestrated ballet of cryptographic operations designed to create confusion and diffusion.&lt;/p&gt;
&lt;p&gt;Let&apos;s break down a single round of DES (remember, there are 16 of these bad boys):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Initial Permutation (IP):&lt;/strong&gt; The 64-bit input block is permuted according to a predefined table. This shuffles the bits around, creating the initial state for the upcoming round.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Transformation:&lt;/strong&gt; A 56-bit key (derived from the original 64-bit key with parity bits removed) is used to generate 16 subkeys, one for each of the 16 rounds. This involves further permutations and shifts.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Splitting:&lt;/strong&gt; The permuted block is split into two 32-bit halves, Left (L) and Right (R).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;F-function (The Heart of DES):&lt;/strong&gt; This is where the magic happens. The R half undergoes a series of operations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Expansion Permutation (E):&lt;/strong&gt; The 32-bit R half is expanded to 48 bits, duplicating some bits according to a specific permutation table. This increases the diffusion effect.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key Mixing:&lt;/strong&gt; The expanded R half is XORed with the round key. This introduces the key&apos;s influence on the encryption process.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;S-box Substitution (The Secret Sauce):&lt;/strong&gt; The 48-bit result is then divided into eight 6-bit chunks. Each chunk is fed into a corresponding S-box (aka Substitution box). Each S-box is a 4x16 lookup table. The 6-bit input selects a row and column in the S-box, and the corresponding 4-bit value is the output. This non-linear substitution is crucial for DES&apos;s security.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Permutation (P):&lt;/strong&gt; The 32-bit output from the S-boxes is permuted again, further scrambling the bits.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;XOR and Swap:&lt;/strong&gt; The output of the F-function is XORed with the L half. Then, the L and R halves are swapped. This completes one round.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Repeat:&lt;/strong&gt; Steps 3-5 are repeated for 16 rounds, using a different subkey for each round.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Final Permutation (IP^-1):&lt;/strong&gt; After 16 rounds, the two halves are combined, and with the resulting 64-bit block a final permutation is carried out. This is an inverse of the initial permutation (IP). This finally produces the ciphertext.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;3DES is simply three applications of DES (encrypt-decrypt-encrypt). It inherits this complex structure but applies it multiple times. Increasing the computational cost but also (theoretically) the security. To be honest, I tried to float 5DES, and 7DES as even more secure versions of 3DES, but the community just rejects them every time.&lt;/p&gt;
&lt;p&gt;This intricate dance of permutations, substitutions, and XOR operations makes DES a seemingly complex yet simple algorithm. The non-linearity introduced by the S-boxes is critical for its security, but the relatively small key size and block size ultimately led to its downfall. Understanding this underlying structure is key (get it!? I just had to!) to acknowledge the limitations of DES which were also the motivations behind the development of more robust ciphers like AES.&lt;/p&gt;
&lt;h2 id=&quot;show-me-the-code&quot;&gt;Show Me The Code!&lt;/h2&gt;
&lt;p&gt;Using DES in production today is about as advisable as wearing a &quot;kick me&quot; post-it note through a high school hallway. But for educational purposes only, here&apos;s a snippet of Python code using the &lt;code&gt;pycryptodome&lt;/code&gt; library. Remember that you &lt;em&gt;shouldn&apos;t&lt;/em&gt; be doing this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;def simplified_f_function(right_half, subkey):
    &quot;&quot;&quot;A simplified F-function (XOR with subkey).&quot;&quot;&quot;
    return right_half ^ subkey

def simplified_des_round(left_half, right_half, subkey):
    &quot;&quot;&quot;A simplified DES round.&quot;&quot;&quot;
    new_right = left_half ^ simplified_f_function(right_half, subkey)
    new_left = right_half
    return new_left, new_right

def simplified_des(plaintext, key, num_rounds=16):
    &quot;&quot;&quot;A simplified DES implementation.&quot;&quot;&quot;
    # Dummy subkey generation (in real DES, this is much more complex)
    subkeys = [key &gt;&gt; i for i in range(num_rounds)]  # Just shifting the key

    # Initial &quot;permutation&quot; (just splitting the plaintext for this example)
    left, right = plaintext[:len(plaintext) // 2], plaintext[len(plaintext) // 2:]

    # Rounds
    for i in range(num_rounds):
        left_int, right_int = int(left, 2), int(right, 2)
        left_int, right_int = simplified_des_round(left_int, right_int, subkeys[i])

        # Convert back to binary strings with padding
        left = bin(left_int)[2:].zfill(len(plaintext) // 2)
        right = bin(right_int)[2:].zfill(len(plaintext) // 2)

    # Final &quot;permutation&quot; (combining the halves)
    ciphertext = left + right
    return ciphertext

# Example usage
plaintext = &quot;1010101001010101&quot;  # Example 16-bit plaintext
key = 0b1111000011001010        # Example 16-bit key (highly insecure)
num_rounds = 4                  # Reduced number of rounds for illustration

ciphertext = simplified_des(plaintext, key, num_rounds)
print(f&quot;Plaintext:  {plaintext}&quot;)   # 1010101001010101
print(f&quot;Ciphertext: {ciphertext}&quot;)  # 1000100000000101101001000011110

# ... and decryption
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Again, this is a museum piece, not even a local-testing-ready implementation.&lt;/p&gt;
&lt;h2 id=&quot;cracking-the-code-from-brute-force-to-sweet-treats&quot;&gt;Cracking the Code: From Brute Force to Sweet Treats&lt;/h2&gt;
&lt;p&gt;The real wake-up call came with the EFF&apos;s Deep Crack project. It demonstrated that brute-forcing a 56-bit key was no longer a theoretical exercise but a practical reality with the computational power available at the time. This, combined with advancements in cryptanalysis, like differential cryptanalysis, further weakened DES&apos;s defenses.&lt;/p&gt;
&lt;p&gt;Wait a minute - what&apos;s this differential cryptanalysis all about? Imagine you&apos;re a codebreaker, and you don&apos;t know the key, but you can encrypt chosen plaintexts. Not so difficult to imagine right? Practically every system today. Differential cryptanalysis exploits the fact that changes in the input (plaintext) lead to predictable differences in the output (ciphertext), even without knowing the key. With this knowledge, by carefully analyzing these differences, or &quot;differentials,&quot; across multiple chosen plaintexts and their corresponding ciphertexts, you can deduce information about the key itself. Neat!&lt;/p&gt;
&lt;p&gt;In the context of DES, differential cryptanalysis targets the S-boxes, those lookup tables that are supposed to be the source of DES&apos;s cryptographic strength are after all a static lookup table. The attack works by identifying pairs of plaintexts with specific differences and observing the resulting differences in the intermediate values after passing through the S-boxes. By analyzing these differential patterns, cryptanalysts can deduce information about the key bits used in each round of DES.&lt;/p&gt;
&lt;p&gt;While differential cryptanalysis doesn&apos;t instantly break DES, it significantly reduces the computational effort required to find the key. It gives you a cheat sheet for a really hard exam – you still need to put in some effort, but you have a significant advantage from having nothing. This type of attack, along with the ever-increasing power of computers, ultimately led to DES&apos;s demise. It also highlighted the need for ciphers with a much stronger resistance to cryptanalysis.&lt;/p&gt;
&lt;p&gt;3DES, while offering a larger key space, wasn&apos;t entirely immune to these types of attacks, although the triple encryption made them significantly more difficult. The Sweet32 attack, exploiting a birthday paradox-style collision vulnerability, demonstrated that even triple encryption could be broken with enough effort (and a whole lot of birthday cakes, apparently). You can delve into the details of this attack &lt;a href=&quot;https://sweet32.info/&quot;&gt;here&lt;/a&gt;. For the official pronouncements, refer to &lt;a href=&quot;https://csrc.nist.gov/news/2023/nist-to-withdraw-sp-800-67-rev-2&quot;&gt;NIST SP 800-67&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;tale-of-the-tape-des-vs-3des-vs-aes&quot;&gt;Tale of the Tape: DES vs. 3DES vs. AES&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;DES&lt;/th&gt;
&lt;th&gt;3DES&lt;/th&gt;
&lt;th&gt;AES&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Key Size&lt;/td&gt;
&lt;td&gt;56 bits&lt;/td&gt;
&lt;td&gt;112/168 bits&lt;/td&gt;
&lt;td&gt;128/192/256 bits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Block Size&lt;/td&gt;
&lt;td&gt;64 bits&lt;/td&gt;
&lt;td&gt;64 bits&lt;/td&gt;
&lt;td&gt;128 bits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security&lt;/td&gt;
&lt;td&gt;Weak&lt;/td&gt;
&lt;td&gt;Deprecated&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Fast (but insecure)&lt;/td&gt;
&lt;td&gt;Slow&lt;/td&gt;
&lt;td&gt;Fast and Secure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mode of Operation&lt;/td&gt;
&lt;td&gt;Limited (ECB is a no-no)&lt;/td&gt;
&lt;td&gt;More flexible, but still...&lt;/td&gt;
&lt;td&gt;Wide range of secure modes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;legacy-lament-the-ghosts-of-crypto-past-and-present-unfortunately&quot;&gt;Legacy Lament: The Ghosts of Crypto Past (and Present, unfortunately)&lt;/h2&gt;
&lt;p&gt;DES and 3DES, despite their known vulnerabilities and exploits, continue to haunt the digital landscape even today. Migrating away from these cryptographic relics can be a herculean task, often fraught with challenges.&lt;/p&gt;
&lt;h2 id=&quot;the-great-crypto-migration-why-its-harder-than-it-sounds&quot;&gt;The Great Crypto Migration: Why It&apos;s Harder Than It Sounds&lt;/h2&gt;
&lt;p&gt;Imagine trying to replace the engine of a car ...that&apos;s moving ...and has a leaking gas line. That&apos;s essentially what migrating away from DES/3DES in a legacy system feels like. Here&apos;s a breakdown of the pain points:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Tight Coupling:&lt;/strong&gt; These old ciphers are often deeply integrated with the application logic, woven into the very fabric of the system. It&apos;s not a simple matter of swapping libraries; it can require extensive code refactoring.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dependency Hell:&lt;/strong&gt; Legacy systems are notorious for their complex and often undocumented dependencies. Changing one thing can trigger a cascade of breakages, turning a simple migration into a descent into the abyss of compatibility issues.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Testing Nightmares:&lt;/strong&gt; Answer this - how do you thoroughly test a cryptographic migration without bringing your entire system to a screeching halt? Comprehensive testing is crucial, but it can be incredibly time-consuming, resource-intensive, and a logistical nightmare, often involving late nights, copious amounts of caffeine, and the occasional existential crisis - at times all at once.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &quot;If It Ain&apos;t Broke...&quot; Mentality:&lt;/strong&gt; Who are we kidding !? Sometimes there&apos;s resistance to change, especially when the system &lt;em&gt;appears&lt;/em&gt; to be functioning (ignoring the ticking time bomb). Convincing stakeholders to invest time, resources, and money in migration can be like pulling teeth.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost:&lt;/strong&gt; All of the above translates to one thing: &lt;em&gt;money&lt;/em&gt;. Migrations are expensive, and justifying the cost to management can be a challenge, especially when there&apos;s no immediate, tangible benefit (besides, you know, ...preventing a potentially catastrophic security breach).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-legacy-lives-on-lessons-for-modern-cryptography&quot;&gt;The Legacy Lives On: Lessons for Modern Cryptography&lt;/h2&gt;
&lt;p&gt;If nothing else, our journey through DES and 3DES reveals a timeless lesson - one about cryptographic evolution. These algorithms demonstrate how even the most trusted security solutions become vulnerable over time. It does a great job at reinforcing the principles of key size importance and algorithm agility that we have discussed in Part 1. As we look ahead to our next deep dive into Blowfish, do remember that today&apos;s gold standard could be tomorrow&apos;s cautionary tale. If you&apos;re still running systems with DES or 3DES, consider this your wake-up call. You only have limited time to start planning that migration to modern algorithms. Your future self (and your security team) will thank you.&lt;/p&gt;
&lt;p&gt;Ready to explore more encryption algorithms? Join me in the next article where we&apos;ll dissect Blowfish and see how it addressed some of DES&apos;s shortcomings while introducing innovations of its own. Do share your thoughts and DES/3DES migration war stories!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[JWT Revocation: Taming the Stateless Beast]]></title><description><![CDATA[Let's just admit it, we've all been seduced by the siren song of stateless authentication with JSON Web Tokens (JWT). No more dealing with…]]></description><link>https://mayankraj.com/blog/jwt-revocation-strategies</link><guid isPermaLink="false">https://mayankraj.com/blog/jwt-revocation-strategies</guid><pubDate>Wed, 01 May 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Let&apos;s just admit it, we&apos;ve all been seduced by the siren song of stateless authentication with JSON Web Tokens (JWT). No more dealing with sticky sessions, no more multiple database lookups for every request… bliss,right? That&apos;s until you need to revoke a token. That&apos;s when it feels less like bliss and more like stepping on a LEGO brick ...barefoot ..at 3 AM. The promise of statelessness, while alluring, introduces a unique challenge no one asked for: how do you effectively revoke access when the very nature of your architecture eschews persistent state?&lt;/p&gt;
&lt;p&gt;In a much more comfortable, traditional, stateful world, revocation is relatively straightforward. You simply invalidate a session on the server, and boom, the user is locked out. But in the ephemeral realm of JWTs, things get a just a little bit trickier. The token is self contained, i.e. it has all the information needed for authentication. This means the server doesn&apos;t inherently &quot;remember&quot; who has access. So, how do we wrestle with this beast of revocation in a stateless world? Let&apos;s dive into the mosh pit of strategies, evaluate them and see what comes out alive.&lt;/p&gt;
&lt;h2 id=&quot;the-usual-suspects-standard-approaches&quot;&gt;The Usual Suspects (Standard Approaches)&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Blacklisting:&lt;/strong&gt; This is the &quot;if it ain&apos;t broke, don&apos;t fix it&quot; approach, albeit with some duct tape and good old-trusty WD-40. You start by maintaining a list of revoked tokens, often in a database or distributed cache. Then you check against it on every single request. Simple in concept, but a potential scalability nightmare. Imagine the latency implications of checking against a list containing thousands of revoked tokens, if not millions. Furthermore, managing this ever-growing list introduces its own set of operational headaches. You also have to flush the &lt;em&gt;expired&lt;/em&gt; token from this list from time to time. All in all, While suitable for smaller applications with a relatively low revocation rates, blacklisting quickly becomes unwieldy in enterprise environments. From a security standpoint, it&apos;s reactive rather than proactive which involves dealing with the aftermath of a potential breach instead of preventing it in real-time. Furthermore, a failure of system leads to token being valid, which means it&apos;s not failsafe in any sizable production environment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Short-lived Tokens:&lt;/strong&gt; This strategy embraces the ephemeral nature of JWTs and cranks it all the way up to eleven. By issuing tokens with very short expiry times (e.g., a few seconds to minutes), the need for explicit revocation diminishes. Compromised tokens become less of a threat because they expire quickly anyway. This time-to-live can be tuned according to the use case and what time seems acceptable for a stray token to be active. However, this introduces the overhead on the systems wherein frequent token refreshes increases the load on your authorization servers. It&apos;s a delicate balancing act, just a tad bit simpler than trying to find the perfect water temperature in a shared shower. Security-wise, it minimizes the window of opportunity for attackers, but requires careful consideration of the refresh token mechanism itself. If you did not guess it already - this mechanism itself becomes a new attack vector if not implemented securely.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;scaling-shenanigans&quot;&gt;Scaling Shenanigans&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Token Introspection:&lt;/strong&gt; Shifting gears from client-side checks to a more predictable server-side validation, token introspection involves querying an authorization server to determine a token&apos;s validity. Instead of maintaining a cumbersome blacklist, the app servers simply ask the authorization server, &quot;Is this token still good?&quot;. This centralizes revocation management while also providing real-time validation. On the flip side, it places a significant burden on the authorization server. Think of it as hiring a highly efficient but also super expensive bouncer for your API party – they&apos;ll keep the riff-raff out, but you&apos;ll need deep pockets to keep them on the payroll. This approach is better suited for scenarios where real-time revocation and centralized control are paramount to everything else.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Refresh Token Rotation:&lt;/strong&gt; This strategy adds a layer of indirection by introducing two sets of tokens - refresh and access tokens. Access tokens remain short-lived, while refresh tokens have a relatively longer lifespan. When an access token expires, the client uses the already issued refresh token to obtain a new access token. This can also be considered as refreshing a token pair, hence &lt;em&gt;refresh token&lt;/em&gt;. More importantly, the old refresh token is invalidated during this process. This sets up mechanisms for indirect revocation i.e. if a refresh token is compromised, only future access is blocked thus limiting the blast radius. It&apos;s like changing the locks after losing your keys – a bit of a hassle, but seems reasonable than having your house ransacked. This approach strikes a good balance between security and usability, while limiting the impact of compromised tokens without constantly interrupting user sessions.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Database Lookup (with Caching):&lt;/strong&gt; While seemingly contradictory to the stateless paradigm we started with, database lookups, when combined with very aggressive caching strategies, can provide a robust and performant revocation solution. Token&apos;s validity is stored in a database, but a distributed cache like Redis acts as a first line of defense, significantly reducing database hits. This approach offers a good balance between security and performance, providing real-time revocation without stressing out your database. Think of it as having a well-organized cheat sheet for your exam – much faster than flipping through the entire textbook and gives you the answers you need. However, this requires careful consideration of cache invalidation strategies to ensure consistency between the cache and the database. The value of still using JWT here lies in the fact that the tokens have other metadata like user details, preferences etc which reduce DB roundtrips.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;code-example-refresh-token-rotation-in-go&quot;&gt;Code (Example: Refresh Token Rotation in Go)&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-go&quot;&gt;// ...Import packages

// Define token claims
type Claims struct {
	UserID uint `json:&quot;user_id&quot;`
	jwt.RegisteredClaims
}

func generateRefreshToken(userID uint, secretKey []byte) (refreshToken string, err error) {
	// Start by creating refresh token
	refreshClaims := &amp;#x26;Claims{
		UserID: userID,
		RegisteredClaims: jwt.RegisteredClaims{
			ExpiresAt: jwt.NewNumericDate(time.Now().Add(7 * 24 * time.Hour)),
			Issuer:    &quot;my_important_issuer&quot;,
			ID:        &quot;unique_id_for_tracking&quot;,
            // ...other claims
		},
	}
	refreshToken, err = jwt.NewWithClaims(jwt.SigningMethodHS256, refreshClaims).SignedString(secretKey)
	return
}

func main() {
	secretKey := []byte(&quot;my_super_secret_key&quot;)
	userID := uint(123)

	refreshToken, err := generateRefreshToken(userID, secretKey)
	if err != nil {
		fmt.Println(&quot;Error generating token:&quot;, err)
		return
	}
	fmt.Println(&quot;Refresh Token:&quot;, refreshToken)

	// ... (Code to store and invalidate refresh tokens would go here)
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;pushing-the-boundaries-emerging-approaches&quot;&gt;Pushing the Boundaries: Emerging Approaches&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Push-based Revocation:&lt;/strong&gt; Cue with WebSockets. Imagine proactively pushing revocation notifications to your services instead of constantly querying from the clients. This minimizes the reliance on checking lists or managing querying endpoints, offering near real-time revocation. The approach is still relatively new and requiring more sophisticated infrastructure, it holds significant promises in terms of enhanced security and performance.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;WebSockets:&lt;/strong&gt; Setup a persistent, bi-directional connection between your authorization server and your resource servers. When a token is revoked, the authorization server pushes a notification over the WebSocket in near-real time. This notification instruct the resource servers to invalidate the token in their local cache. This approach provides the most responsive revocation mechanism, significantly reducing the window of vulnerability. However, it requires managing persistent connections, a cache on the resource servers and also handling of potential disruptions.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Server-Sent Events (SSE):&lt;/strong&gt; We don&apos;t really need a bi-directional communication channel and this is where SSE comes in. It&apos;s a lighter-weight alternative to WebSockets. SSE allows the authorization server to push updates to resource servers. This is ideal for scenarios where real-time revocation is desired, but the complexity of WebSockets is not called for. SSE provides a unidirectional communication channel, making it simpler to implement, maintain and debug.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Challenges:&lt;/strong&gt; Implementing a push-based revocation introduces complexities related to scalability as well as reliability. Your messaging infrastructure should be able to handle the volume of revocation events, and resource servers need to process these events efficiently. Robust error handling and retry mechanisms are essential to account for the inevitable network interruptions. While not a trivial undertaking, the potential benefits in terms of responsiveness and reduced latency make this an area worth exploring for applications demanding high-security and high-performance applications. Imagine the possibilities!&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-good-the-bad-and-the-ugly-trade-offs--&quot;&gt;The Good, the Bad, and the Ugly (Trade-offs) -&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Pros&lt;/th&gt;
&lt;th&gt;Cons&lt;/th&gt;
&lt;th&gt;Security Implications&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blacklisting&lt;/td&gt;
&lt;td&gt;Simple to implement&lt;/td&gt;
&lt;td&gt;Doesn&apos;t scale well, potential latency issues&lt;/td&gt;
&lt;td&gt;Reactive, relies on identifying compromised tokens after potential misuse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Short-lived Tokens&lt;/td&gt;
&lt;td&gt;Reduces impact of compromised tokens&lt;/td&gt;
&lt;td&gt;Increased token refresh overhead&lt;/td&gt;
&lt;td&gt;Minimizes attack window but doesn&apos;t eliminate risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token Introspection&lt;/td&gt;
&lt;td&gt;Centralized revocation management&lt;/td&gt;
&lt;td&gt;Increased load on authorization server&lt;/td&gt;
&lt;td&gt;Real-time validation, reduces risk of stale data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refresh Token Rotation&lt;/td&gt;
&lt;td&gt;Limits damage from compromised refresh tokens&lt;/td&gt;
&lt;td&gt;Requires careful management of refresh tokens&lt;/td&gt;
&lt;td&gt;Good balance, limits damage without constant interruptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Database Lookup (with caching)&lt;/td&gt;
&lt;td&gt;Balance of security and performance&lt;/td&gt;
&lt;td&gt;Relies on caching infrastructure&lt;/td&gt;
&lt;td&gt;Effective with proper cache invalidation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;choosing-the-right-strategy-a-balancing-act&quot;&gt;Choosing the Right Strategy: A Balancing Act&lt;/h2&gt;
&lt;p&gt;Choosing the right JWT revocation strategy is not a decision for the faint of heart. It requires carefully balancing security requirements, performance goals, all while staying within the budget constraints. Consider the following factors:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scale:&lt;/strong&gt; How many users and tokens will your system handle?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security Requirements:&lt;/strong&gt; How sensitive is the data you are protecting?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Performance Goals:&lt;/strong&gt; What are your acceptable latency targets?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Budget:&lt;/strong&gt; How much can you invest in revocation infrastructure?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Development Resources:&lt;/strong&gt; What expertise and resources do you have available?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For smaller applications with less stringent security requirements to deal with, short-lived tokens or blacklisting might be enough. For much larger, enterprise-level applications with high-security needs, token introspection, refresh token rotation, or database lookup with caching are more suitable. Push-based revocation is an emerging approach which is gaining popularity because it offers promising performance benefits but requires more specialized expertise and infrastructure.&lt;/p&gt;
&lt;h2 id=&quot;bottom-line&quot;&gt;Bottom Line&lt;/h2&gt;
&lt;p&gt;Sadly but to no one&apos;s surprise - there&apos;s no silver bullet for JWT revocation in stateless environments. The best approach depends on a nuanced understanding of your specific needs and scale of operations. Carefully consider the trade-offs, balancing security requirements, performance goals, and budget constraints. Make sure that you get creative and combine approaches – a hybrid strategy might be the perfect fit. And remember, a well-designed revocation strategy can be the difference between a secure system and a security nightmare.&lt;/p&gt;
&lt;p&gt;What&apos;s your go-to strategy? Share your thoughts and experiences. Let&apos;s learn from each other. I&apos;m especially interested in hearing from folks who&apos;ve implemented push-based revocation in production!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[JWTs: Beyond the Buzzwords - What You Really Need to Know]]></title><description><![CDATA[Let's talk Json Web Tokens aka JWT, shall we? Not just a surface level skim, but a much needed deep dive. You've probably already worked…]]></description><link>https://mayankraj.com/blog/jwt-deep-dive-beyond-basics</link><guid isPermaLink="false">https://mayankraj.com/blog/jwt-deep-dive-beyond-basics</guid><pubDate>Tue, 16 Apr 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Let&apos;s talk Json Web Tokens aka JWT, shall we? Not just a surface level skim, but a much needed deep dive. You&apos;ve probably already worked with them, and maybe even fought a JWT-fueled fire drill it the past. But are we really doing our bit in squeezing all the juice out of these little JSON nuggets? Let me tell you something ...let me tell you something - there&apos;s a whole damn universe inside a JWT just waiting to be explored and we are going to do just that!&lt;/p&gt;
&lt;h2 id=&quot;decoding-the-damn-thing&quot;&gt;Decoding the Damn Thing&lt;/h2&gt;
&lt;p&gt;Think of a JWT like a self-contained, digitally signed and verifyable envelope. It&apos;s all about secure, portable data for your applications. Three parts, each separated by dots: Header, Payload, and the Signature. Sounds simple enough, right? Well, to no one&apos;s surprise the devil&apos;s in the details. Let&apos;s peel back the layers and explore them.&lt;/p&gt;
&lt;h2 id=&quot;anatomy-of-a-jwt&quot;&gt;Anatomy of a JWT&lt;/h2&gt;
&lt;p&gt;Here’s a JWT broken down, piece by piece. We&apos;ll get into what each part &lt;em&gt;really&lt;/em&gt; means:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;// Header (Base64Url encoded)
{
  &quot;alg&quot;: &quot;HS256&quot;,       // Specifies the algorithm
  &quot;typ&quot;: &quot;JWT&quot;          // Just letting everyone know it&apos;s a JWT - standard practice
}

.

// Payload (Base64Url encoded) - This good stuff
{
  &quot;sub&quot;: &quot;user123&quot;,     // Subject - who this token is about
  &quot;name&quot;: &quot;John Doe&quot;,   // User&apos;s name - or whatever you want to store
  &quot;exp&quot;: 1678886400,    // Expiration time - crucial for security
  &quot;roles&quot;: [&quot;admin&quot;, &quot;editor&quot;] // User roles - for authorization magic
}

.

// Signature (Base64Url encoded) - The cryptographic seal of approval
HMACSHA256(
  base64UrlEncode(header) + &quot;.&quot; +
  base64UrlEncode(payload),
  your-256-bit-secret // Keep this REALLY secret - like your Netflix password
)
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;decoding-a-real-jwt&quot;&gt;Decoding a Real JWT&lt;/h2&gt;
&lt;p&gt;Enough of the theory! Let&apos;s get practical. Here&apos;s an actual encoded JWT as it would appear in the wild (don&apos;t worry, the secret&apos;s fake):&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkpvaG4gRG9lIiwiaWF0IjoxNTE2MjM5MDIyfQ.SflKxwRJSMeKKF2QT4fwpMeJf36POk6yJV_adQssw5c
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You can decode this using quite a few standard library like &lt;code&gt;jsonwebtoken&lt;/code&gt; in Node.js (version 16.x.x as of this writing):&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;const jwt = require(&quot;jsonwebtoken&quot;);

const token =
  &quot;eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiIxMjM0NTY3ODkwIiwibmFtZSI6IkpvaG4gRG9lIiwiaWF0IjoxNTE2MjM5MDIyfQ.SflKxwRJSMeKKF2QT4fwpMeJf36POk6yJV_adQssw5c&quot;;

try {
  const decoded = jwt.verify(token, &quot;your-256-bit-secret&quot;);
  console.log(decoded);
} catch (err) {
  console.error(&quot;Invalid token:&quot;, err);
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;signing-ceremony-choosing-your-cryptographic-weapon-wisely&quot;&gt;Signing Ceremony: Choosing Your Cryptographic Weapon Wisely&lt;/h2&gt;
&lt;p&gt;The beauty of JWT is in the signing algorithm. It is very critical and not the area where you slack off. It’s one of assets in your digital fortress against token forgery. HS256 (HMAC with SHA-256) is relatively fast while also using a shared secret. This is great for internal systems where you control both the ends. But if that secret ever makes it&apos;s way out of the walled garden, you&apos;re toast. As an alternative, RS256 (with SHA-256) makes use of a public-private keyspair. It bring in much stronger security and thus making it ideal for external APIs as well as distributed systems. This however comes with a performance cost of working with RSA.&lt;/p&gt;
&lt;p&gt;Choosing the right algorithm is always a balancing act – security versus speed. There&apos;s never a one-size-fits-all, especially when the topic at hand is security. You have to consider your applications context in the larger ecosystem, your exposure, the application&apos;s blast radius, the target audience and their willingness to invest into breaking your mechanisms. Ask yourself - are you dealing with highly sensitive data? What is your performance budget? Are you ready to spend the next 7 weekends looking at signatures?&lt;/p&gt;
&lt;h2 id=&quot;jwt-superpowers-beyond-just-simple-authentication&quot;&gt;JWT Superpowers: Beyond Just Simple Authentication&lt;/h2&gt;
&lt;p&gt;While JWTs are often just used as digital bouncers checking IDs at the door but they are much much more than just that. Fine-grained authorization? Let&apos;s go! You can simply embed roles and permissions directly in the payload. Skip all that redundant database trips. Have a need for single sign-on (SSO)? JWTs can make that happen without even a sweat. Letting users slide between services without much friction. Do I hear - Auditing? logging? Inspect theJWT claims to track user activity like a hawk.&lt;/p&gt;
&lt;h2 id=&quot;performance-anxiety-taming-jwts-at-scale&quot;&gt;Performance Anxiety: Taming JWTs at Scale&lt;/h2&gt;
&lt;p&gt;I hate to do this but there&apos;s bad news as well - JWTs do end up adding overhead and this overhead is not negligible. At scale, that overhead can become a real pain in the app. Keeping your payloads concise and to the bare minimum can make a really big difference. Think about how frugally you pack for a trip - only the essentials and not even a napkin more. Every extra byte eventually adds up. Signature verification is computationally expensive, and so caching them appropriately can be your friend. With that now you have an additional caching layer to deal with. And remember, while statelessness has it&apos;s own perks, it also makes revocation a lot more tricky. Short expiration times and a blacklist cache at the host of more computational power can help you sleep at night, that is if at all in the first place.&lt;/p&gt;
&lt;h2 id=&quot;jwt-vs-the-world&quot;&gt;JWT vs. The World&lt;/h2&gt;
&lt;p&gt;JWTs, while fun aren’t the only name in town. Opaque tokens are much simpler to use and revocable by design but require server-side storage. We also have the good old session cookies which are easy to manage but can struggle with scalability in any size of distributed workloads. How can we forget the API keys which while great for service-to-service communication, lacks the granular control tha today&apos;s complex app need. In the end, choosing the right authentication mechanism depends on your application&apos;s specific needs and the enviornment that it&apos;ll operate in. It&apos;s always agood idea to not follow the hype. Take the day off, get a cup of cooffee and think it through.&lt;/p&gt;
&lt;h2 id=&quot;jwt-in-the-wild-tales-from-the-trenches&quot;&gt;JWT in the Wild: Tales from the Trenches&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Case Study 1: The Exploding Payload:&lt;/strong&gt; I once inherited a service that crammed &lt;em&gt;everything&lt;/em&gt; under the sun into the JWT payloads - think user data, preferences, the kitchen sink. While it worked great initially, even giving a false sense of security that the developers have reduced the DB roundtrips and with that improved the users experience. But as the tale goes - the application scaled and these very massive tokens crippled performance. Network calls became sluggish, users complained, costs skyrocketed. Lesson learned: keep payloads lean and mean!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pitfall 1: The &quot;none&quot; Algorithm Nightmare:&lt;/strong&gt; Let me be very clear here - never, ever, ever ever, ever ever ever ever even think of using the &quot;none&quot; algorithm! Not even when &quot;testing on the local&quot;, definately not on the &quot;test env&quot;. It disables signature verification altogether, making your JWTs as secure as a screen door on a submarine. While a rookie mistake, I’ve seen it happen in real customer facing apps. Don&apos;t be that person for a change.&lt;/p&gt;
&lt;h2 id=&quot;wielding-jwts-effectively&quot;&gt;Wielding JWTs Effectively&lt;/h2&gt;
&lt;p&gt;JWTs are powerful tools - no questions on that. But they require some careful considerations. Understand their strengths and weaknesses, and only then choose the right signing algorithm for your specific security and performance needs. Always be mindful of performance, especially at scale. When used correctly, they can enhance your application&apos;s security and user experience. Used incorrectly, they can be a source of not only headaches but also the even more expensive - vulnerabilities. So choose wisely, my friend.&lt;/p&gt;
&lt;p&gt;What are your experiences with JWTs, especially at scale? Do share them! Let&apos;s learn from each other.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Insights in Plaintext: Deep Dives into Cryptographic Algorithms]]></title><description><![CDATA[Gear up! We're about to embark on an exciting journey through the intricate maze that is cryptography! This is Part 1 of our "Insights in…]]></description><link>https://mayankraj.com/blog/insights-in-plaintext-introduction</link><guid isPermaLink="false">https://mayankraj.com/blog/insights-in-plaintext-introduction</guid><pubDate>Mon, 01 Apr 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Gear up! We&apos;re about to embark on an exciting journey through the intricate maze that is cryptography! This is Part 1 of our &quot;Insights in Plaintext&quot; series. In here where we&apos;ll strip away the complexity and them dive deep into the algorithms that keep our digital world secure. My goal is to explore these algorithms with you. I&apos;m just curious about how our data stays safe online. Hopefully this series will arm you with practical knowledge about encryption - from battle-tested classics to quantum-ready innovations. Let&apos;s kick things off by going over the fundamentals and setting the stage for our deep dives into specific algorithms in upcoming articles.&lt;/p&gt;
&lt;h2 id=&quot;the-wild-west-of-data-security&quot;&gt;The Wild West of Data Security&lt;/h2&gt;
&lt;p&gt;Let&apos;s be honest for a second, the digital world is a wild west jungle. Data breaches, identity theft, online snooping – it&apos;s a freakin&apos; free-for-all out there. In 2023 over 32% of digital attacks were intended for data theft and leak (&lt;a href=&quot;https://www.ibm.com/reports/threat-intelligence&quot;&gt;Source: IBM Security X-Force Threat Intelligence Index 2023&lt;/a&gt;). That&apos;s why encryption is no longer a nice-to-have line item, but a damn necessity. Consider it your data&apos;s very own personal bodyguard, decked out in Kevlar and ready to rumble at a moment&apos;s notice. From your grandma&apos;s cat videos to your company&apos;s top-secret cola recipe, everything has value to someone, and encryption is the key to keeping it in safe hands ...or hard drives.&lt;/p&gt;
&lt;h2 id=&quot;encryption-101-ciphers-keys-and-the-quantum-threat&quot;&gt;Encryption 101: Ciphers, Keys, and the Quantum Threat&lt;/h2&gt;
&lt;p&gt;At its core, encryption essentially scrambles your data using a secret code (a cipher) and a key, much like a lockbox. The cipher is the locking mechanism, and the key is what opens it. Only those with the correct key can unscramble the data. We have two main types:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Symmetric Encryption:&lt;/strong&gt; Uses the same key for both encryption and decryption. While it&apos;s fast and efficient, key security is paramount. Examples include AES (Advanced Encryption Standard), the current gold standard, and older algorithms like DES and 3DES.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Asymmetric Encryption:&lt;/strong&gt; Uses two keys: one for encryption (a public key) and one for decryption (a private key). It&apos;s like a mailbox with a public slot and a private key to control who can open it. While it&apos;s slower in practice it&apos;s incredibly secure for key exchange and digital signatures. RSA (Rivest–Shamir–Adleman) and ECC (Elliptic Curve Cryptography) are prime examples.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then there&apos;s &lt;strong&gt;Hashing&lt;/strong&gt;, a one-way function transforming data into a unique, fixed-size string (a hash). This is crucial for verifying data integrity. Any change in the data alters the hash in a significant manner. SHA-3 (Secure Hash Algorithm 3) is a leading hashing algorithm today. HMAC (Hash-Based Message Authentication Code) uses hashing for authentication.&lt;/p&gt;
&lt;p&gt;Now, a word of caution. While these methods are strong today, quantum computing is just around the corner. Quantum computers, with their immense processing power, could potentially break widely used encryption algorithms at play today. The good news? The industry is developing quantum-resistant cryptography. We&apos;ll discuss this more later in the series.&lt;/p&gt;
&lt;h2 id=&quot;decrypting-the-future-whats-coming-next&quot;&gt;Decrypting the Future: What&apos;s Coming Next&lt;/h2&gt;
&lt;p&gt;This series will arm you with the knowledge you need to navigate the complex landscape of encryption. Here&apos;s a sneak peek at what&apos;s coming up:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;DES/3DES: The Golden Oldies (and Why They&apos;re Not So Golden Anymore):&lt;/strong&gt; We&apos;ll explore these once-dominant algorithms, and examine their drawbacks in today&apos;s age.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Blowfish: A Fast Fish in a Big Pond:&lt;/strong&gt; Discover the strengths and limitations of this speedy algorithm and its suitability for modern applications.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HMAC: Hashing for Authentication:&lt;/strong&gt; Learn how HMAC leverages hashing to provide robust authentication while also providing data integrity checks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SHA-3: The Hashing Heavyweight Champion:&lt;/strong&gt; We&apos;ll delve into the low-level intricacies of SHA-3 and the significant role it plays in securing digital signatures as well as data integrity.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Diffie-Hellman: The Key Exchange Mastermind:&lt;/strong&gt; Understand how Diffie-Hellman enables secure key exchange, even over insecure channels.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RSA: The Asymmetric Anchor:&lt;/strong&gt; Explore the enduring power of RSA and its continued importance in secure communication.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AES: The Reigning King of Symmetric Encryption:&lt;/strong&gt; We&apos;ll dissect AES, the current gold standard, and understand why it&apos;s so widely adopted.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ECC: The Elliptic Curve Enigma:&lt;/strong&gt; Discover the elegance and efficiency of ECC, especially in resource-constrained environments like mobile devices.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OTP: The One-Time Wonder:&lt;/strong&gt; Explore the theoretical un-breakability of OTP and the practical challenges that limit its widespread use.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Get ready to unlock the secrets of encryption!&lt;/p&gt;
&lt;h2 id=&quot;the-road-ahead-your-journey-into-cryptography&quot;&gt;The Road Ahead: Your Journey into Cryptography&lt;/h2&gt;
&lt;p&gt;We&apos;ve covered enough ground for today even if it&apos;s just the fundamental principles of encryption and the looming challenges of quantum computing. I hope that this foundation will serve as our launching pad for deep dives into specific algorithms in upcoming articles. Whether it&apos;s about securing sensitive data, building secure applications, or just being fascinated by cryptography - understanding these concepts is crucial in our increasingly digital world.&lt;/p&gt;
&lt;p&gt;Ready to continue this cryptographic journey? Do share your questions and thoughts. Let&apos;s decode these challenges together!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Great Cloud Egress Escape: Why Your Cloud Bill Gives You Nightmares]]></title><description><![CDATA[So, you've migrated to the cloud - let's say AWS, and you're basking in the glow of scalability and cost savings... right? Well not so soon…]]></description><link>https://mayankraj.com/blog/egress-cost-optimisations</link><guid isPermaLink="false">https://mayankraj.com/blog/egress-cost-optimisations</guid><pubDate>Sat, 09 Mar 2024 06:30:00 GMT</pubDate><content:encoded>&lt;p&gt;So, you&apos;ve migrated to the cloud - let&apos;s say AWS, and you&apos;re basking in the glow of scalability and cost savings... right? Well not so soon my friend. The next bill arrives, and it&apos;s higher than Taylor Swift&apos;s backstage pass. What gives? One sneaky culprit often overlooked is the gold old data egress fees. The cost of transferring data &lt;em&gt;out&lt;/em&gt; of the cloud. It&apos;s a trap - you were warned. If not, I&apos;ll take this opportunity to discuss egress with you...&lt;/p&gt;
&lt;h2 id=&quot;the-egress-elephant-in-the-room&quot;&gt;The Egress Elephant in the Room&lt;/h2&gt;
&lt;p&gt;Data transfer &lt;em&gt;within&lt;/em&gt; a cloud region (say an AWS region) is usually free (or very very cheap), lulling you into a false sense of security. But the moment your data needs to get out of that cozy AWS region walled garden – whether to your on-premise servers, another cloud provider, or even just your users browsing the internet – BAM! Egress charges hit you like nothing has ever before.&lt;/p&gt;
&lt;p&gt;For the remainder of this post, I&apos;ll use AWS as an example to drive a point - but be rest assured that it&apos;s not alone in the game.&lt;/p&gt;
&lt;h2 id=&quot;decoding-the-data-egress-matrix&quot;&gt;Decoding the Data Egress Matrix&lt;/h2&gt;
&lt;p&gt;AWS&apos;s data egress pricing is nothing short of a choose-your-own-adventure novel, except all of the endings involve you paying. The costs vary based on many factors like...&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Destination:&lt;/strong&gt; Transferring data to the world-wide-internet is generally a bit cheaper than transferring to your own data center. Go figure that out...&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Volume:&lt;/strong&gt; The more data you move ...well ...the more you pay (duh). But on the other end of the spectrum volume discounts &lt;em&gt;do&lt;/em&gt; exist, so bulk transfers can be slightly less painful.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Region:&lt;/strong&gt; Egress charges differ between different AWS regions. This means choosing the right region for your workload can leave you with some savings in the end.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Service:&lt;/strong&gt; To no one&apos;s surprise - different AWS services have their own egress pricing quirks. S3, for example, has its own set of rules. Fun times all around.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;show-me-the-money-or-rather-how-much-its-going-to-cost-me&quot;&gt;Show Me the Money! (Or Rather, How Much It&apos;s Going to Cost Me)&lt;/h2&gt;
&lt;p&gt;Let&apos;s shift gears and get down to brass tacks. Imagine that you need to transfer over 1TB of data. Here’s a &lt;em&gt;hypothetical&lt;/em&gt; cost comparison (because AWS pricing changes faster than crypto prices):&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Egress Destination&lt;/th&gt;
&lt;th&gt;&lt;em&gt;Hypothetical&lt;/em&gt; Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Internet&lt;/td&gt;
&lt;td&gt;$90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same Region&lt;/td&gt;
&lt;td&gt;$0 (or very minimal)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Different Region&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Direct Connect (to your data center)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;em&gt;Disclaimer: Please note that these figures are for illustrative purposes only. The actual costs will vary based on your specific setup, usage, region, timing, etc.&lt;/em&gt; You certainly don&apos;t want to be surprised by a bill in the end that looks nothing short of the national debt.&lt;/p&gt;
&lt;h2 id=&quot;static-assets-the-cdn-vs-the-rest&quot;&gt;Static Assets: The CDN vs. The Rest&lt;/h2&gt;
&lt;p&gt;Now let&apos;s consider a fairly common scenario - a website serving ~1TB of static assets (images, JavaScript, CSS) per month. Let&apos;s compare three approaches available to us:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Serving assets directly from an EC2 instance:&lt;/strong&gt; This option is nothing less than the egress cost equivalent of setting your money on fire. Every asset - image, script, stylesheet and more, downloaded incurs data transfer charges. Let&apos;s say, hypothetically, transferring 1TB out of &lt;code&gt;us-east-1&lt;/code&gt; costs around $90. To add to it, you will need a beefy instance running in the first place. We&apos;ll ignore that for this comparison.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Serving via API Gateway:&lt;/strong&gt; Slightly better, but still far from ideal. API Gateway adds its own costs over and above the egress fees. You&apos;re paying for the data transfer and then some more for API Gateway to process the data. Let&apos;s assume this adds another $10-$20 to our bill, bringing the total to $100-$110.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Serving via CloudFront (aka CDN):&lt;/strong&gt; This is the clear winner! CloudFront not only caches your static assets at edge locations around the world but also reduces the load on your origin. When a user requests for an asset, it&apos;s served from the closest edge location. This minimizes data transfer from your origin server (S3 or EC2) and dramatically reduces the egress costs. This also reduces the load on your servers thus allowing you to scale them down - further reducing the bill. Think of it as having a wide network of squirrels stashing copies of your website all over the garden. With CloudFront, you might pay significantly less for the same data transfer out to the internet. This could be in the range of ~$25 for that same 1TB, and even less for requests served from the cache.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In this scenario, using CloudFront helped you save $65-$85 per month compared to serving directly from EC2. Not to even mention that you save further by reducing the capacity needed to serve the same traffic. Finally as your traffic grows, those savings multiply. Cha-ching!&lt;/p&gt;
&lt;h2 id=&quot;supercharging-your-cdn-tips-and-tricks&quot;&gt;Supercharging Your CDN: Tips and Tricks&lt;/h2&gt;
&lt;p&gt;Okay CDN sounds convincing and now you&apos;re using a CDN like CloudFront. Amazing! But are you making sure to squeeze every last drop of available performance and cost savings out of it? Don&apos;t scratch your head, here are a few quick tips:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Caching Policies to Control Your Destiny:&lt;/strong&gt; Caching policies end up deciding how long assets are stored at edge locations. Configure these policies wisely based on your requirements:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Example 1: Aggressive Caching for Static Assets:&lt;/strong&gt; Assets like your brand logo - rarely change. Hopefully. Cache it for a long long time, like a year. This means every request for the logo will be served from the CDN&apos;s cache itself. This means hitting your origin server (and in turn incurring egress fees) only once a year. Neat!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Example 2: Moderate Caching for Frequently Updated Content:&lt;/strong&gt; Let&apos;s say you have a popular blog online (maybe like this one!?). While new articles are posted regularly, they might not be posted every minute. Cache these articles for a relatively shorter duration, perhaps a few hours or a day. This balances freshness with cache efficiency. An update to the post will be visible only after the Cache expires. While you&apos;ll incur some egress costs when new articles are published, eventually the requests will be served from the cache itself.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Example 3: No Caching for Highly Dynamic Content:&lt;/strong&gt; For data that changes frequently, like I don&apos;t know ...Mr. Elon Musk&apos;s rant on a topic - caching might not be appropriate. In such cases, requests should always hit your origin server. But thankfully, your users will at least have up-to-the-second data!&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;HTTPS - Security with Savings:&lt;/strong&gt; Two-for-one !? Serving content over HTTPS goes hand in hand for security, but it can also end up impacting your CDN costs. CloudFront for example, offers different pricing tiers for secure traffic. Make sure you calculate with the pricing structure in mind and only then choose the most cost-effective option. At times, using dedicated IPs or different SSL certificate configurations can leave you to discover some surprising cost differences.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Geo-Restrictions: Keep Your Content Where It Belongs:&lt;/strong&gt; If all your users are in specific geographic regions then it makes sense to use geo-restrictions. This will prevent access from other areas and thus reduce unnecessary data transfer as well as egress costs. Not only this but it can also improve performance for your target audience.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;the-development-delusion-egress-what-egress&quot;&gt;The Development Delusion: &quot;Egress? What Egress?&quot;&lt;/h2&gt;
&lt;p&gt;We&apos;ve all been there - bright-eyed and bushy-tailed in the early stages of development, focused on building cool features and shipping code. Cost optimization? Pfft, we&apos;ll worry about that later, right? Wrong!! Ignoring egress costs during the initial designing stage is like ignoring a leaky faucet – it seems small at first, but it can quickly turn into a massive flood of expenses. When you&apos;re prototyping and testing - sure, the data transfer will seem negligible. But as your application scales, the user base grows, the traffic increases, and with that those egress costs can snowball faster than a runaway snowball rolling down a really big, snowy hill.&lt;/p&gt;
&lt;h2 id=&quot;taming-the-egress-beast-strategies-for-savings&quot;&gt;Taming the Egress Beast: Strategies for Savings&lt;/h2&gt;
&lt;p&gt;So in the end - it does matter. How do you wrestle these costs into submission? Here are a few nifty tactics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Minimize Data Transfer:&lt;/strong&gt; Seems obvious, but often overlooked. Do you really need to move all that data? Can you not process it within AWS first and only then move the results out? Think lean, think mean, think data-transfer-reduction machine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Content Delivery Networks (CDNs):&lt;/strong&gt; For static content like images, videos etc, a CDN like CloudFront can be your best friend. CDNs cache your content closer to users thus reducing the amount of data that needs to travel from your origin server. This in turn reduces egress costs, all while reducing the load on your servers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Compression:&lt;/strong&gt; Do you want another no-brainer? How about compressing your data before transferring it? You&apos;d be surprised by how many folks forget this basic step. We all pack our suitcases efficiently right? You can fit more in, it&apos;s easier to carry and cheaper to transport.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Right-Sizing Instances:&lt;/strong&gt; Abundant choices but choosing the right instance type for your workload can minimize unnecessary data transfer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use Specialised Tools:&lt;/strong&gt; Explore specialized tools like AWS DataSync for efficient data transfer and cost optimization. Someone has thought about it, why not use it?&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;show-me-the-code-because-every-good-blog-post-needs-code&quot;&gt;Show Me the Code! (Because Every Good Blog Post Needs Code)&lt;/h2&gt;
&lt;p&gt;Here&apos;s a high-level sample Python snippet demonstrating a simple compression technique using &lt;code&gt;gzip&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import gzip
import shutil

def compress_file(input_file, output_file):
    with open(input_file, &apos;rb&apos;) as f_in:
        with gzip.open(output_file, &apos;wb&apos;) as f_out:  # Gzipping like a boss
 shutil.copyfileobj(f_in, f_out)


# Example usage (because we&apos;re helpful like that)
compress_file(&apos;massive_data_file.txt&apos;, &apos;massive_data_file.txt.gz&apos;)  # Boom! Shrinkage!
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;dont-let-egress-eat-your-lunch&quot;&gt;Don&apos;t Let Egress Eat Your Lunch&lt;/h2&gt;
&lt;p&gt;Data egress costs can be a real pain-in-the-a**. But with a little ahead-of-time planning and some clever strategies, you can tame the beast. Don&apos;t let those sneaky egress fees turn your cloud dream into a financial nightmare. What are your favorite egress-busting strategies? Do share them!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Bits, Bytes, and Buttons: The Unexpected But Obvious Link Between Elevators and Hard Drives]]></title><description><![CDATA[Ever wondered about the connection between elevators and hard drives? Well I have, and surprisingly while on an elevator! If you give it a…]]></description><link>https://mayankraj.com/blog/scheduling-in-elevator-harddrive</link><guid isPermaLink="false">https://mayankraj.com/blog/scheduling-in-elevator-harddrive</guid><pubDate>Fri, 02 Feb 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Ever wondered about the connection between elevators and hard drives? Well I have, and surprisingly while on an elevator! If you give it a thought - both involve scheduling, resource management, and optimization. Both have many pending request to go from A-B, for one it&apos;s floors, and for the other the disk location.&lt;/p&gt;
&lt;h2 id=&quot;up-and-down-and-all-around-scheduling-in-distributed-systems&quot;&gt;Up and Down and All Around: Scheduling in Distributed Systems&lt;/h2&gt;
&lt;p&gt;Elevators and hard drives have to handle multiple requests all while minimizing wait times and seek times. Both make use of scheduling algorithms in order to optimize performance. Here are few such commonly used algorithm.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SCAN:&lt;/strong&gt; The pointer moves in a single direction, servicing requests, then turns around and moves back. Like an elevator going to the top most floor, then back down to the ground floor. It&apos;ll wait on any floor in between which have been requested.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LOOK:&lt;/strong&gt; Similar to SCAN, but the head only travels as far as the last request in each direction. The elevator will still strictly travel in one direction, but will reverse it&apos;s direction of there is no request for any other floor in that direction.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SSTF (Shortest Seek Time First):&lt;/strong&gt; Prioritizes the nearest request and thus minimizing seek time but at the cost of starving distant requests. In this case, the elevator may end up only serving the first few floors of a 50 floor building.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;C-SCAN (Circular SCAN):&lt;/strong&gt; Moves in one direction, returning to the beginning after reaching the end. In this case, the elevator only goes up, and when it reaches the top it starts over again from the ground floor going up.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here&apos;s a simplified Python implementation of these algorithms:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;# algos.py
import matplotlib.pyplot as plt

def scan(requests, head, direction, disk_size=200):
    requests = sorted(set(requests))  # Remove duplicates and sort
    left = [r for r in requests if r &amp;#x3C; head]
    right = [r for r in requests if r &gt;= head]

    if direction == &quot;left&quot;:
        result = left[::-1] + [0] + right
    else:  # direction == &quot;right&quot;
        result = right + [disk_size - 1] + left[::-1]

    print(f&quot;SCAN: {result}&quot;)
    return result

def look(requests, head, direction):
    requests = sorted(set(requests))  # Remove duplicates and sort
    left = [r for r in requests if r &amp;#x3C; head]
    right = [r for r in requests if r &gt;= head]

    if direction == &quot;left&quot;:
        result = left[::-1] + right
    else:  # direction == &quot;right&quot;
        result = right + left[::-1]

    print(f&quot;LOOK: {result}&quot;)
    return result

def sstf(requests, head):
    requests = list(set(requests))  # Remove duplicates
    result = []
    current_head = head
    while requests:
        closest_request = min(requests, key=lambda x: abs(x - current_head))
        result.append(closest_request)
        current_head = closest_request
        requests.remove(closest_request)

    print(f&quot;SSTF: {result}&quot;)
    return result

def cscan(requests, head, direction, disk_size=200):
    requests = sorted(set(requests))  # Remove duplicates and sort
    right = [r for r in requests if r &gt;= head]
    left = [r for r in requests if r &amp;#x3C; head]

    result = right + [disk_size - 1] + [0] + left

    print(f&quot;C-SCAN: {result}&quot;)
    return result

def calculate_total_seek_time(schedule, start_position):
    total_seek_time = 0
    current_position = start_position

    for request in schedule:
        seek_distance = abs(request - current_position)
        total_seek_time += seek_distance
        current_position = request

    return total_seek_time

def plot_seek_operations(algorithm_name, schedule, start_position):
    plt.figure(figsize=(10, 6))
    plt.title(f&quot;{algorithm_name} Disk Scheduling&quot;)
    plt.xlabel(&quot;Request Sequence&quot;)
    plt.ylabel(&quot;Disk Position&quot;)

    x = range(len(schedule) + 1)
    y = [start_position] + schedule

    plt.plot(x, y, marker=&apos;o&apos;)
    plt.grid(True)
    plt.show()

# Example usage:
if __name__ == &quot;__main__&quot;:
    requests = [176, 79, 34, 60, 92, 11, 41, 114]
    head = 50
    direction = &quot;right&quot;
    disk_size = 200

    print(&quot;Initial state:&quot;)
    print(f&quot;Requests: {requests}&quot;)
    print(f&quot;Initial head position: {head}&quot;)
    print(f&quot;Direction: {direction}&quot;)
    print(f&quot;Disk size: {disk_size}&quot;)
    print()

    for algorithm in [scan, look, sstf, cscan]:
        if algorithm == sstf:
            result = algorithm(requests.copy(), head)
        elif algorithm == look:
            result = algorithm(requests.copy(), head, direction)
        else:
            result = algorithm(requests.copy(), head, direction, disk_size)

        total_seek_time = calculate_total_seek_time(result, head)
        print(f&quot;Total seek time: {total_seek_time}&quot;)
        plot_seek_operations(algorithm.__name__.upper(), result, head)
        print()
&lt;/code&gt;&lt;/pre&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;$&gt; python3 algos.py
    Initial state:
    Requests: [176, 79, 34, 60, 92, 11, 41, 114]
    Initial head position: 50
    Direction: right
    Disk size: 200

    SCAN: [60, 79, 92, 114, 176, 199, 41, 34, 11]
    Total seek time: 337

    LOOK: [60, 79, 92, 114, 176, 41, 34, 11]
    Total seek time: 291

    SSTF: [41, 34, 11, 60, 79, 92, 114, 176]
    Total seek time: 204

    C-SCAN: [60, 79, 92, 114, 176, 199, 0, 11, 34, 41]
    Total seek time: 389
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;In a distributed system it becomes tricky to coordinate multiple &quot;elevators&quot; (or nodes). We end up introducing challenges like fault tolerance, consistency and conflict resolution. Algorithms like Paxos or Raft, which rely on distributed consensus often use a leader node to manage the scheduling.&lt;/p&gt;
&lt;h2 id=&quot;the-enterprise-elevator-scaling-up-our-metaphor&quot;&gt;The Enterprise Elevator: Scaling Up Our Metaphor&lt;/h2&gt;
&lt;p&gt;With the basics down, let&apos;s crank up this metaphor up to 11 to see how it applies to enterprise-scale systems. Because let&apos;s be real for a second, if you&apos;re writing code, you&apos;re not just managing one elevator in a small four-story building. You are instead running the entire transportation system of not one but five megacities.&lt;/p&gt;
&lt;h3 id=&quot;distributed-elevators-a-cluster-of-confusion&quot;&gt;Distributed Elevators: A Cluster of Confusion&lt;/h3&gt;
&lt;p&gt;Imagine that you are tasked with optimizing the elevator system for the gold standard - Burj Khalifa. Well all of a sudden our simple scheduling algorithms look about as effective as a kids paper airplane in a hurricane. We&apos;re talking over 163 floors, 57 elevators, and thousands of people trying to get to their destinations before the coffee cools down.&lt;/p&gt;
&lt;p&gt;This is where our elevator-hard drive analogy really starts to shine, especially in the context of distributed systems. Let&apos;s break it down:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Load Balancing&lt;/strong&gt;: Just like you would make sure to distribute requests across multiple hard drives in a RAID array, you need to efficiently distribute passengers across multiple elevators. The goal isn&apos;t just to reduce the wait time, that the by-product - the goal is to preventing system failures and ensuring consistent performance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Predictive Analytics&lt;/strong&gt;: Just like the coffee machine next to you, modern elevator systems use &lt;em&gt;AI&lt;/em&gt; to predict the usage patterns. Sound familiar eh? It&apos;s the same principle behind predictive read-ahead available in modern SSDs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Fault Tolerance&lt;/strong&gt;: What happens when an elevator breaks down? After you have rescued the people - you need to also redistribute the remaining load. This is just like a distributed database handling node failures on the fly.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Here&apos;s a simplified example of how we might model this in code:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import random
from collections import deque

class Elevator:
    def __init__(self, id, max_floor):
        self.id = id
        self.current_floor = 1
        self.destination = 1
        self.passengers = []
        self.max_floor = max_floor
        self.direction = &quot;up&quot;
        self.status = &quot;operational&quot;

    def move(self):
        if self.current_floor &amp;#x3C; self.destination:
            self.current_floor += 1
            self.direction = &quot;up&quot;
        elif self.current_floor &gt; self.destination:
            self.current_floor -= 1
            self.direction = &quot;down&quot;

    def add_passenger(self, passenger):
        self.passengers.append(passenger)
        self.update_destination()

    def update_destination(self):
        if self.passengers:
            self.destination = max(p.destination for p in self.passengers)
        else:
            self.destination = self.current_floor

class ElevatorSystem:
    def __init__(self, num_elevators, max_floor):
        self.elevators = [Elevator(i, max_floor) for i in range(num_elevators)]
        self.max_floor = max_floor
        self.requests = deque()

    def request_elevator(self, from_floor, to_floor):
        # In a real system, we&apos;d use a more sophisticated algorithm here
        available_elevators = [e for e in self.elevators if e.status == &quot;operational&quot;]
        if available_elevators:
            elevator = min(available_elevators, key=lambda e: abs(e.current_floor - from_floor))
            elevator.add_passenger(Passenger(from_floor, to_floor))
        else:
            self.requests.append((from_floor, to_floor))

    def simulate_step(self):
        for elevator in self.elevators:
            if elevator.status == &quot;operational&quot;:
                elevator.move()
                # Handle passenger drop-offs, etc.

        # Handle queued requests if elevators become available
        while self.requests and any(e.status == &quot;operational&quot; for e in self.elevators):
            from_floor, to_floor = self.requests.popleft()
            self.request_elevator(from_floor, to_floor)

class Passenger:
    def __init__(self, start_floor, destination):
        self.start_floor = start_floor
        self.destination = destination

# Usage
system = ElevatorSystem(num_elevators=57, max_floor=163)
for _ in range(1000):  # Simulate 1000 steps
    system.simulate_step()
    if random.random() &amp;#x3C; 0.1:  # 10% chance of a new request each step
        from_floor = random.randint(1, system.max_floor)
        to_floor = random.randint(1, system.max_floor)
        system.request_elevator(from_floor, to_floor)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Obviously this code is a gross oversimplification, but it illustrates the point we are after: managing a distributed system, whether it&apos;s elevators or databases, requires designing around resource allocation, fault tolerance, and optimizing. All of this for the global efficiency rather than local optimums.&lt;/p&gt;
&lt;h2 id=&quot;elevating-your-thinking&quot;&gt;Elevating Your Thinking&lt;/h2&gt;
&lt;p&gt;So, what&apos;s the point of this you ask? It&apos;s not just to make you think twice next time you&apos;re in an elevator (though that&apos;s a fun side effect). The real takeaway is about how we approach a seemingly artifical topic of system design at scale, with something that we have a day to day experience of:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Adapt and Predict&lt;/strong&gt;: A well designed system, whether it&apos;s moving people or data - anticipate needs, adapt to challenge and falls back to the next optimal move.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Think Holistically&lt;/strong&gt;: Just as an efficient elevator system considers the entire building, it&apos;s layout, where the elevators are places and the high traffic for the floor with the candy shop, we need to design our systems with the big picture in mind as well.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Balance Efficiency and Fairness&lt;/strong&gt;: Pure efficiency (like prioritising serving the closest request) can lead to unfairness at large. Your design needs to balance everything with at sometimes competing objectives.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Prepare to Failure&lt;/strong&gt;: In any sufficiently complex system, something will go wrong - and eventualy everything will go wrong. Design with redundancy and graceful degradation in mind from the get go.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Optimize, Observe, Optimize&lt;/strong&gt;: The job is never done! Just as the smart elevator systems are constantly learning and adjusting themselfs, our data systems should be continuously optimized based on the real-world usage patterns.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Remember that whether you&apos;re moving bits or bodies, the principles of efficient, scalable system design remain the same. Next time you&apos;re architecting a distributed system - pause and take a moment to think about elevators. Why? Because it might just elevate your solution to new heights.&lt;/p&gt;
&lt;p&gt;Now, if you&apos;ll excuse me, I still have several floors of data to go through...&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Base64 Encoding Demystified: Padding, Security, and Practical Applications]]></title><description><![CDATA[Let's look at Base64, shall we? There's a good change that you have already seen it atleast twice today. From URLs to email attachments…]]></description><link>https://mayankraj.com/blog/base64-padding-and-security</link><guid isPermaLink="false">https://mayankraj.com/blog/base64-padding-and-security</guid><pubDate>Thu, 11 Jan 2024 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Let&apos;s look at Base64, shall we? There&apos;s a good change that you have already seen it atleast twice today. From URLs to email attachments, they are everywhere. But what exactly is this funny looking encoding scheme? Why does it sometimes end with those equal signs (&lt;code&gt;==&lt;/code&gt;)? Why isn’t it encryption? Why was my application to Standford rejected? Well answer to some lies in my introspection, and for others let&apos;s dive into this essential piece of the web&apos;s infrastructure.&lt;/p&gt;
&lt;h2 id=&quot;base64-in-the-world-wild-web&quot;&gt;Base64 in the World-Wild-Web&lt;/h2&gt;
&lt;p&gt;Before we get into the technical nitty gritty, let&apos;s look at where exactly does Base64 earns its spot in the enterprise:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Image Data URIs:&lt;/strong&gt; Used to encode and embedding small images directly in HTML or CSS using. This can improve page load times. Think icons or small logos, with no aditional requests.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Email Attachments:&lt;/strong&gt; Encoding attachments like PDFs, images, etc. as Base64 strings allows us club them together within the email email body. This stramlines email handling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data Storage and Retrieval:&lt;/strong&gt; Base64 makes storing binary data in text-based databases or configuration files more possible. This means the data lies closely together with each other.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security Tokens and Credentials:&lt;/strong&gt; While Base64 isn&apos;t encryption, you’ll often see it used within the security contexts. For example in encoding parts of JSON Web Tokens (JWTs). But remember - encoding is &lt;em&gt;not&lt;/em&gt; encryption!&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;decoding-the-encoding-how-base64-works&quot;&gt;Decoding the Encoding: How Base64 Works&lt;/h2&gt;
&lt;p&gt;Base64 at their core encodes binary data using only printable ASCII characters. This is very important for interoperability because some systems struggle with raw binary data. They can potentially misinterpret byte sequences as control characters and cause unexpected behavior. Base64 provides a safe, text-based encoding to avoid these issues in the first place.&lt;/p&gt;
&lt;p&gt;Here’s the process at a very high level - binary data is chopped into 6-bit groups, each mapped to a specific printable ASCII character using a lookup table (A-Z, a-z, 0-9, +, /).&lt;/p&gt;
&lt;h2 id=&quot;visualizing-the-transformation-mayankrajcom-goes-base64&quot;&gt;Visualizing the Transformation: &quot;MayankRaj.com&quot; Goes Base64&lt;/h2&gt;
&lt;p&gt;Let&apos;s see this in action with &quot;MayankRaj.com&quot;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;# Step 1: String to Bytes (ASCII)
MayankRaj.com  -&gt;  77 97 121 97 110 107 82 97 106 46 99 111 109

# Step 2: Bytes to Binary (8-bit)
77  -&gt; 01001101
97  -&gt; 01100001
121 -&gt; 01111001
97  -&gt; 01100001
110 -&gt; 01101110
107 -&gt; 01101011
82  -&gt; 01010010
97  -&gt; 01100001
106 -&gt; 01101010
46  -&gt; 00101110
99  -&gt; 01100011
111 -&gt; 01101111
109 -&gt; 01101101

# Step 3: Regrouping into 6-bit Chunks (with annotations)
010011 010110 000101 111001 011000 010110  &amp;#x3C;- First six 6-bit groups
111001 101011 010100 100110 000101 101010  &amp;#x3C;- Next six 6-bit groups
001011 100110 001101 101111 011011 010000  &amp;#x3C;- Last six 6-bit groups (padded)


# Step 4: Mapping to Base64 Characters (Tabular Format)
| 6-bit Chunk | Base64 Character |
|---|---|
| 010011 | M |
| 010110 | W |
| 000101 | F |
| 111001 | 5 |
| 011000 | Y |
| 010110 | W |
| 111001 | 5 |
| 101011 | r |
| 010100 | R |
| 100110 | a |
| 000101 | F |
| 101010 | q |
| 001011 | L |
| 100110 | a |
| 001101 | M |
| 101111 | v |
| 011011 | b |
| 010000 | Q |


# Final Base64 Encoded String:
TWF5YW5rUmFqLmNvbQ== (The == indicates padding)

&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;show-me-the-action&quot;&gt;Show Me the Action!&lt;/h2&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import base64

binary_data = b&apos;Hello, world!&apos;
encoded_data = base64.b64encode(binary_data)
print(encoded_data)  # Output: b&apos;SGVsbG8sIHdvcmxkIQ==&apos;

decoded_data = base64.b64decode(encoded_data)
print(decoded_data)  # Output: b&apos;Hello, world!&apos;
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;the-security-angle---encoding-vs-encryption&quot;&gt;The Security angle - Encoding vs. Encryption&lt;/h2&gt;
&lt;p&gt;Let&apos;s be very sure of one key aspect of Base64 - It is &lt;em&gt;encoding&lt;/em&gt;, not &lt;em&gt;encryption&lt;/em&gt;. Encoding just changes the data&apos;s representation, but not its underlying meaning. Anyone can decode it, let alone a computer. On the flip side, Encryption scrambles the data in such a way that you can only make sense of it with a key. This key is usually called encryption key and is handled with cared. Make sure to not confuse the two! Use proper encryption for security, and encoding for reliable data transfer.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[The Port Party: A Crash Course on Ephemeral Ports and Their Security Implications]]></title><description><![CDATA[Let's continue our deep dive into a critical, yet often overlooked aspect of distributed systems — ephemeral ports. We looked at what they…]]></description><link>https://mayankraj.com/blog/ephemeral-ports-in-the-network-stack</link><guid isPermaLink="false">https://mayankraj.com/blog/ephemeral-ports-in-the-network-stack</guid><pubDate>Sun, 03 Dec 2023 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Let&apos;s continue our deep dive into a critical, yet often overlooked aspect of distributed systems — ephemeral ports. We looked at what they are with &lt;a href=&quot;/blog/ephemeral-ports/&quot;&gt;Journey Through the Silicon&lt;/a&gt; and also looked at proxies with &lt;a href=&quot;/blog/forward-reverse-proxy-explained/&quot;&gt;From Forward to Reverse Proxies&lt;/a&gt;. Let&apos;s look at how these two play with each other. Ephemeral ports are always in the backstage control room, but trust me, you don&apos;t want to ignore them. They might be small, but they can cause big problems if you don&apos;t understand how they work.&lt;/p&gt;
&lt;p&gt;Think of ephemeral ports as the temporary phone numbers used by your client applications to communicate with the servers. Each time any client initiates a connection, the operating system assigns it a port number randomly picked from a reserved range. These ports are tightly coupled with this connection and only active during the connection lifetime — hence the name &quot;ephemeral&quot;. When the job is done, the port is put back in the bowl to be picked up for another connection.&lt;/p&gt;
&lt;p&gt;So, why should you care? After all these are temporary, ever-changing ports often short lived for the lifecycle of the connection. Well - Because they play a crucial role in making sure a smooth communication in distributed systems can take place. They are the reason for the connection to happen in the first place. You wouldn&apos;t want your shiny new fancy distributed application to become a tangled mess, right?&lt;/p&gt;
&lt;h3 id=&quot;load-balancers-who-let-the-port-numbers-out-&quot;&gt;Load Balancers: Who Let the Port-Numbers Out !?&lt;/h3&gt;
&lt;p&gt;Let&apos;s say you have a bunch of servers, which host your backend. These are the army powering your distributed application. How do you direct traffic to them efficiently? Cue the load balancer, the overworked traffic cop that takes all the incoming connections and splits them evenly across your servers.&lt;/p&gt;
&lt;p&gt;Load balancers are typically deployed in front of your backend services as the first line of defence. Today load balancers can also operate at Layer 7, which means they not only distribute the incoming traffic, but have the ability to peek inside the request. This also means they have to terminate the incoming request from the client, and establish a new one with your servers.&lt;/p&gt;
&lt;p&gt;When a client request comes in, the load balancer assigns it a unique ephemeral port at its level. It then directs it to an appropriate backend server. It&apos;s like the load balancer gives each client a temporary phone number, and then connects them to the right server based on its internal routing logic.&lt;/p&gt;
&lt;p&gt;Load balancers, a type of proxies was discussed in much more details over at &lt;a href=&quot;/blog/forward-reverse-proxy-explained/&quot;&gt;From Forward to Reverse Proxies&lt;/a&gt;.&lt;/p&gt;
&lt;h4 id=&quot;sticky-sessions-the-double-edged-sword-of-performance&quot;&gt;Sticky Sessions: The Double-Edged Sword of Performance&lt;/h4&gt;
&lt;p&gt;Applications are not always built &quot;ideal&quot;. Some end up having session state stored locally on the server. To tackle this one common strategy used by load balancers is &quot;sticky sessions.&quot; It makes sure to keep a client connected to the same server throughout a series of requests. This can improve performance by utilizing cached data which is tied to a given server. Imagine a shopping cart: just because developers choose to use a load balancer, you don&apos;t want to start adding items on one server and then end up on a different one at checkout, right? That would be a recipe for a frustrated customer (and potentially a lost sale)!&lt;/p&gt;
&lt;p&gt;To flip the coin around, sticky sessions can cause issues too. If the server that a given client&apos;s request is attached to goes down, the load balancer needs to find a new server and new port. This could potentially disrupt the client&apos;s ongoing session. To add to it, if you have a slow server, all the clients stuck on it will experience performance problems as sticky sessions will not allow them to go to another server. It&apos;s like having the party guest stuck in a slow moving elevator - everyone is late to the party, and with that the fun is getting disrupted.&lt;/p&gt;
&lt;h3 id=&quot;nat-gateways-translating-traffic-like-a-pro-with-some-ephemeral-magic&quot;&gt;NAT Gateways: Translating Traffic Like a Pro (With Some Ephemeral Magic)&lt;/h3&gt;
&lt;p&gt;Ever heard of NAT? No !? It stands for Network Address Translation. It&apos;s a common means for hiding internal IP addresses from the wild outside world. NAT gateways behave like translation services - they are constantly converting the client&apos;s IP address and port to different ones, so that your internal servers can communicate safely. Above all, all of this without exposing themselves directly to the clients.&lt;/p&gt;
&lt;p&gt;Usually NAT gateways utilize SNAT (Source Network Address Translation) to rewrite the source IP address and port. Think of it as the gateway taking your client&apos;s phone number, switching it up, and then giving it to the server. Generally only the NAT has these mapping with it. This helps protect your internal network while maintaining communication flow. The ports are allocated at NAT level, for either end of the communication.&lt;/p&gt;
&lt;h4 id=&quot;the-ephemeral-port-pool-a-shared-resource&quot;&gt;The Ephemeral Port Pool: A Shared Resource&lt;/h4&gt;
&lt;p&gt;But just like life - there&apos;s a catch. NAT gateways often use ephemeral ports for this translation process. This ends up creating an ephemeral port pool which is then shared across all translated connections. What this means is that when your gateway is handling a ton of traffic, it can run into issues with port exhaustion. This in turn could lead to dropped connections or performance degradation. It&apos;s like everyone wants Sheldon&apos;s spot because it&apos;s ..well ..the perfect spot - it&apos;s a recipe for chaos!&lt;/p&gt;
&lt;h4 id=&quot;tuning-nat-gateways-for-optimal-performance&quot;&gt;Tuning NAT Gateways for Optimal Performance&lt;/h4&gt;
&lt;p&gt;Just like with Linux iptables, NAT gateways can be tweaked as well which can help improve their performance. The key here is to configure the appropriate values for ephemeral ports, connection timeout, etc which eventually can lead to optimizing the translation process. All of this will significantly impact your NAT&apos;s ability to scale and respond. NAT should be a part of the stack that are not noticed by anyone. You should be able to set it and forget it. So time spent in optimisations here are well worth it.&lt;/p&gt;
&lt;h3 id=&quot;proxies-middleman-or-bottleneck-how-ephemeral-ports-behave-in-proxy-scenarios&quot;&gt;Proxies: Middleman or Bottleneck? How Ephemeral Ports Behave in Proxy Scenarios&lt;/h3&gt;
&lt;p&gt;Proxies act as intermediaries between the clients and servers. When set as forward proxies they are forwarding requests from clients to servers. When as reverse proxies they end up as a gateway to your backend services. It&apos;s like having a friendly bartender who takes orders from guests and delivers them to the kitchen. It&apos;s forwarding requests from one point-of-view, and proxying request from the others.&lt;/p&gt;
&lt;p&gt;When a client makes a request via a proxy, the proxy will almost certainly assign an ephemeral port to manage this connection. What this means is that the proxy becomes a temporary relay station, handling the traffic flowing between the client and the server. On the other hand, if the proxy isn&apos;t configured appropriately, it can create a bottleneck, especially when dealing with a high volume of requests. This is where tuning knobs like HTTP keep-alive, connection pooling etc come into play. They can not only help reduce the burden on the proxy but also improve efficiency. You always choose to have your Amazon packages be delivered in &quot;fewer boxes&quot; right ? Why not do the same with your network packages?&lt;/p&gt;
&lt;h4 id=&quot;proxy-technologies-a-choice-of-tools&quot;&gt;Proxy Technologies: A Choice of Tools&lt;/h4&gt;
&lt;p&gt;Choices are everywhere, and proxies are not behind. Popular technologies like Squid, Varnish Cache, and HAProxy are often used in large scale distributed systems. These offer different features and optimizations like connection pooling, caching, load balancing and more. All of these can be critical for managing ephemeral ports effectively. And as we have seen time and again - managing ephemeral ports are key to seamless operations.&lt;/p&gt;
&lt;h3 id=&quot;scaling-shenanigans-how-ephemeral-ports-affect-distributed-system-performance&quot;&gt;Scaling Shenanigans: How Ephemeral Ports Affect Distributed System Performance&lt;/h3&gt;
&lt;p&gt;In the world of distributed systems, scaling is the buzzword. And ephemeral ports play a crucial role in achieving high performance and scalability. But they can also end up being the source of unexpected performance issues, particularly in a high-throughput environments.&lt;/p&gt;
&lt;h4 id=&quot;real-world-examples-port-exhaustion-in-action&quot;&gt;Real-World Examples: Port Exhaustion in Action&lt;/h4&gt;
&lt;p&gt;Consider a large online website handling a massive influx of traffic during a concert ticket sale. You have been there, haven&apos;t you? The requests come in all-at-once, quite literally. The system might experience port exhaustion due to the overwhelming number of connections being established via the load balancers, NAT gateways and proxies. This can lead to connections being dropped, throttling and overall customer frustration.&lt;/p&gt;
&lt;h4 id=&quot;performance-tuning-tips-avoiding-port-starvation&quot;&gt;Performance Tuning Tips: Avoiding Port Starvation&lt;/h4&gt;
&lt;p&gt;In my experience the best way to avoid port exhaustion, is to consider the following...&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Increase Ephemeral Port Range:&lt;/strong&gt; Expand the range of ports available for ephemeral use. Every system has a certain number of ports, and when you know that the server would need a high number of port - it&apos;s best to allocate it ahead of time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use Connection Pooling:&lt;/strong&gt; Tackle the problem at the source - reduce the number of connections established in the first place. Reusing existing connections wherever possible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Setup Load Balancing:&lt;/strong&gt; Optimize the load balancing algorithms. If the load balancer can make sure all the available resources are used optimally, you won&apos;t have to scale any one individually or overallocate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Monitor Port Usage:&lt;/strong&gt; Track ephemeral port usage, so that you can identify potential bottlenecks ahead of time and address them. Remember - logs, logs and even more logs.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;security-considerations-guarding-your-port-party&quot;&gt;Security Considerations: Guarding Your Port Party&lt;/h3&gt;
&lt;p&gt;Ephemeral ports facilitate communication but they also pose a major security challenges. If a bad actor knows that ephermal ports will be used for the requests, it could be worthwhile to check &quot;all&quot; the ports - maybe once in a while the bad actor will hit gold! Ephemeral ports can be vulnerable to various attacks like port scanning, port exhaustion attacks, and brute force attacks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firewall Rules:&lt;/strong&gt; Firewall rules can be used to limit access to specific ports. Thus restricting connections from unauthorized sources.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Network Segmentation:&lt;/strong&gt; Breaking up the network into smaller, isolated segments help prevent malicious attacks from spreading. This also reduces the blast radius.&lt;/p&gt;
&lt;h3 id=&quot;at-last---avoiding-ephemeral-port-hell&quot;&gt;At last - Avoiding Ephemeral Port Hell&lt;/h3&gt;
&lt;p&gt;Ephemeral ports are crucial components of the modern distributed system stack. Better understanding how they work within components like load balancers, NAT gateways and proxies is the key to building reliable and scalable applications. Avoiding common pitfalls like port exhaustion and properly configuring your infrastructure, can ensure that your system operates as it is designed to.&lt;/p&gt;
&lt;p&gt;Remember, ephemeral ports might be small behind the scenes players, but they have a big impact on your overall performance. So, be aware of them when designing your applications. Don&apos;t let those pesky little ephemeral ports turn your system into a party gone wrong!&lt;/p&gt;</content:encoded></item><item><title><![CDATA[From Forward to Reverse Proxies: Enhancing Network Performance and Security]]></title><description><![CDATA[From Forward to Reverse Proxies: Enhancing Network Performance and Security Let's be honest, folks. Any engineer who has worked for more…]]></description><link>https://mayankraj.com/blog/forward-reverse-proxy-explained</link><guid isPermaLink="false">https://mayankraj.com/blog/forward-reverse-proxy-explained</guid><pubDate>Fri, 10 Nov 2023 18:30:00 GMT</pubDate><content:encoded>&lt;h2 id=&quot;from-forward-to-reverse-proxies-enhancing-network-performance-and-security&quot;&gt;From Forward to Reverse Proxies: Enhancing Network Performance and Security&lt;/h2&gt;
&lt;p&gt;Let&apos;s be honest, folks. Any engineer who has worked for more than a year has already seen it all – from spaghetti code that would make a seasoned Italian chef weep to production outages that would make even a grown engineer cry. And let&apos;s not even talk about those &quot;super-critical-path-breaking-urgent&quot; requests from marketing that magically appear at 4:59 PM on a Friday.&lt;/p&gt;
&lt;p&gt;But here&apos;s the thing – we don&apos;t have to navigate this crazy world alone. Proxies, my friends, are the unsung heroes of scalable and a more stable architecture, always there to lend a helping hand (or at least a helpful network hop).&lt;/p&gt;
&lt;h3 id=&quot;forward-vs-reverse-proxies-two-sides-of-the-same-coin-sort-of&quot;&gt;Forward vs. Reverse Proxies: Two Sides of the Same Coin (Sort of)&lt;/h3&gt;
&lt;p&gt;Forget the technical definitions just for a second. A proxy, in its simplest form, is just a middle-person. It&apos;s that friend who always intercepts your calls when your ex is blowing up your phone, except instead of drama, it&apos;s dealing with network traffic.&lt;/p&gt;
&lt;p&gt;Now, there are two main flavors of proxies that you will come across – the forward and reverse.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Forward Proxies: The Client&apos;s Best Friend (and Firewall Foe)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For a second picture this: you&apos;re at a tech conference, desperately trying to update your LinkedIn profile before that crucial networking event. As with every conference that you have been to, the conference Wi-Fi is not reliable. With that now you&apos;re stuck behind a firewall.&lt;/p&gt;
&lt;p&gt;That&apos;s where a &lt;strong&gt;forward proxy&lt;/strong&gt; comes in handy. It&apos;s like having a secret tunnel right under the firewall, allowing you to bypass those pesky restrictions and access the websites you need (even if it&apos;s just to stalk your new &quot;acquaintance&quot; on LinkedIn... I mean, research potential employers).&lt;/p&gt;
&lt;p&gt;Forward proxies are masters of disguise, masking your IP address and making it look like all your requests are coming from the proxy server itself. It&apos;s like browsing the web incognito, but with a slightly more sophisticated fedora and with that trusted hacker-sunglasses combo.&lt;/p&gt;
&lt;p&gt;But wait, there&apos;s more! Forward proxies can also cache content, just like your browser does, but on a much larger scale. This means faster load times for everyone and less strain on the origin server (which is probably already sweating under the pressure of a thousand engineers trying to update their LinkedIn profiles simultaneously).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Popular forward proxies:&lt;/strong&gt; Squid, Apache HTTP Server (with mod_proxy), HAProxy&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Reverse Proxies: Shielding Your Servers from the Wild, Wild Web&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If forward proxies are all about protecting the client, &lt;strong&gt;reverse proxies&lt;/strong&gt; are the guardians of your servers. Think of them as the bouncers at the club – all requests go through them first, and only the legit ones get through to the party inside (your precious servers).&lt;/p&gt;
&lt;p&gt;Reverse proxies are the multi-tasking superheroes of the server world. It is designed to distribute traffic like a seasoned air traffic controller, preventing any single server from getting overwhelmed (and crashing and burning in a fiery blaze of server errors). They can handle the SSL encryption, offloading the heavy lifting from your servers and making sure your users&apos; data is safe and sound. And yes, they can even cache content, because who doesn&apos;t love a good performance boost?&lt;/p&gt;
&lt;p&gt;That&apos;s not all, here&apos;s where things get really interesting – reverse proxies can also be used for some pretty amazing stuff. Stuff like A/B testing and blue-green deployments. We can instruct this air traffic controller to direct traffic to different set of servers, all while your users continue to watch the add that they should be watching, and definitely not blocking off with an ad-blocker (...right?).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Popular Reverse Proxies:&lt;/strong&gt; Nginx, HAProxy, Varnish Cache&lt;/p&gt;
&lt;h3 id=&quot;lets-talk-business-show-me-the-code&quot;&gt;Let&apos;s talk business: Show Me The Code!&lt;/h3&gt;
&lt;p&gt;Alright, enough with the metaphors. Let&apos;s see some real-world action with one of the popular options out there - Nginx, the Swiss Army Knife of web servers and proxies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scenario 1: Load Balancing&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Imagine you just launched a killer new feature that is bound to change the course of humanity, and your website traffic is exploding faster than a bag of microwave popcorn. Nginx can swoop in and distribute those requests across multiple servers like a boss, preventing any single server from becoming a smoldering pile of silicon.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-nginx&quot;&gt;http {
    upstream myapp {
        server web1.example.com;   # Server 1: Probably sipping margaritas somewhere
        server web2.example.com;   # Server 2: Trying to keep up with the requests
        server web3.example.com;   # Server 3: Wishing it had invested in more RAM
    }

    server {
        listen 80;
        server_name www.example.com;

        location / {
            proxy_pass http://myapp;  # Nginx doing its magic, redirecting traffic
        }
    }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Scenario 2: SSL Termination – Because Security Shouldn&apos;t Give Your Servers a Nervous Breakdown&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;SSL encryption is like the secret handshake of the internet, ensuring that only authorized parties can access sensitive data. But handling all that encryption and decryption can be a real drag on your servers, especially when they&apos;re already juggling a million other tasks. These involve relatively complex cryptographic operations, which are not easy to pull off.&lt;/p&gt;
&lt;p&gt;Nginx to the rescue! Again !? YES ! It can handle SSL termination at the proxy level, freeing up your servers to focus on what they do best – serving up those beautiful web pages with even more adds.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-nginx&quot;&gt;server {
    listen 443 ssl;
    server_name www.example.com;

    ssl_certificate /path/to/certificate.crt;   # Your shiny SSL certificate
    ssl_certificate_key /path/to/private.key;   # Don&apos;t lose this!

    location / {
        proxy_pass http://localhost:8080;  # Assuming your web server is running locally
    }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h3 id=&quot;proxies-in-the-wild&quot;&gt;Proxies in the Wild&lt;/h3&gt;
&lt;p&gt;Here&apos;s the thing about modern software development – just like the real world, it&apos;s messy. We&apos;re talking microservices, containers, serverless functions, microservices built with serverless functions, microservices, microservices built with serverless functions hosted on containers... you get the point. But in the heart of this chaotic jungle, proxies are our trusted guides, helping us navigate the complexities and emerge victorious (...or maybe at least with just our sanity intact).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Load Balancing: Not Just for Your Overworked Web Servers&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Remember when load balancing was as simple as distributing traffic across a few web servers? Wait what !? You do !!?? Those were the good old days, my friend. Now we&apos;re dealing with sprawling microservices architectures, where dozens, hundreds, or even &lt;em&gt;hundreds of thousands&lt;/em&gt; of services are fighting for resources.&lt;/p&gt;
&lt;p&gt;But fear not, for proxies are here to save the day (again!). They&apos;ve evolved and leveled up their load balancing game, acting as intelligent traffic cops at different layers of our architecture. They route requests, optimize resource utilization, and prevent those dreaded cascading failures that can bring down your entire system faster than you can say &quot;blame it on the network.&quot;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Service Mesh&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Ah, the service mesh – the latest and arguably the greatest buzzword in the world of microservices. It&apos;s like the glue that holds your entire distributed system together, ensuring that all those tiny services can communicate and collaborate without descending into an inevitable anarchy.&lt;/p&gt;
&lt;p&gt;And guess what plays a starring role in this intricate dance of microservices? You guessed it – proxies!&lt;/p&gt;
&lt;p&gt;In a service mesh, proxies come out to be the heroes, quietly working behind the scenes to handle service discovery, routing, security, and even observability - all while maintaining their pazzaz.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Istio:&lt;/strong&gt; Built on the ever-popular Envoy proxy, Istio is like the Swiss Army Knife of service meshes. It&apos;s packed with features, but be warned – it comes with a very very very very steep learning curve.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Linkerd:&lt;/strong&gt; If Istio is the overachieving older sibling, Linkerd is the arguably cool, minimalist younger one. It&apos;s known for its simplicity, performance, and ease of use – perfect for dipping your toes into the service mesh pool without getting overwhelmed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consul Connect:&lt;/strong&gt; HashiCorp knows their stuff when it comes to infrastructure management. Consul Connect is no exception in that list. It integrates seamlessly with their service discovery tool, making it a breeze to set up, scale and manage.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Global Traffic Routing: Because Distance Shouldn&apos;t Matter&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;As your user base expands beyond your city, state, or even continent, you need to ensure that everyone gets a fast and reliable experience, regardless of their geographical location. And guess what can help you achieve that? You got it once again – proxies! This is getting too predictable now isn&apos;t it ?&lt;/p&gt;
&lt;p&gt;Proxies are the masters of global traffic routing, using a variety of techniques to bring your application closer to your users:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CDN (Content Delivery Network):&lt;/strong&gt; CDNs are like having a network of caches strategically placed around the world which themself act as an extention to your servers. They not only store copies of your static content (images, videos, etc.) closer to your users, so they don&apos;t have to wait an eternity for things to load but also reduce the load on your origin servers by not asking it for content time and again.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Geo-DNS:&lt;/strong&gt; Remember that time you accidentally booked a flight to the wrong London? (Been there - done that) Geo-DNS is like the GPS of the internet, routing users to the nearest data center - all of that based on just their IP address.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anycast Routing:&lt;/strong&gt; Anycast routing is like having a team of identical twins working for you – it sends traffic to the closest available server, ensuring high availability and fault tolerance.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;the-challenges-of-managing-proxies-at-scale-because-with-great-power-comes-great-responsibility-and-headaches&quot;&gt;The Challenges of Managing Proxies at Scale: Because With Great Power Comes Great Responsibility (...and Headaches)&lt;/h3&gt;
&lt;p&gt;Let&apos;s take a step back, and reflect on a developers life shall we - Life can be like a puzzle—at first, it&apos;s fun. Then, it’s like trying to finish it with quite a few missing pieces. Scaling proxies? That’s doing the puzzle while the pieces change shape in your hands.&lt;/p&gt;
&lt;p&gt;Here are a few challenges you might encounter on your proxy-powered journey:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Configuration Management:&lt;/strong&gt; With dozens or even hundreds of proxies scattered across your infrastructure, keeping track of all those configurations can make your head spin. You may now need a proxy layer to communicate the configurations to all your proxies, which themself are in the proxy layer. Centralized configuration management tools and automation are your best friends here (trust me, you don&apos;t want to be manually updating configurations at 3 AM ...on saturday night ...after things are on fire).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Observability:&lt;/strong&gt; How do you know if your proxies are actually doing their job? You need eyes everywhere! We all want more logs, and there&apos;s not enough of it. Robust monitoring, logging, and tracing are essential to better understand traffic flow, identifying bottlenecks, all while troubleshooting issues in a proxy-heavy environment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Performance Tuning:&lt;/strong&gt; Remember that extra hop we talked about? Well, it can come at a cost if you&apos;re not careful about it. Proxies can introduce latency if not configured correctly, and sometimes even if configured properly. You need to optimize those caching mechanisms, connection pooling settings, and all those other knobs and dials that make proxies sing.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;so-you-think-you-want-to-be-a-proxy-master&quot;&gt;So, You Think You Want to Be a Proxy Master?&lt;/h3&gt;
&lt;p&gt;This is where you come in.. your time to rise and shine. Have you wrestled with proxies in your own projects? What war stories can you share about those late-night debugging sessions with a once trusted proxy that now a rogue culprit? Do share your experiences, tips, and questions.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Journey Through the Silicon : Network communications via Ephemeral Ports]]></title><description><![CDATA[Do you ever stop to think about the tiny marvels of technology that make our digital world go 'round? We're diving into the building blocks…]]></description><link>https://mayankraj.com/blog/ephemeral-ports</link><guid isPermaLink="false">https://mayankraj.com/blog/ephemeral-ports</guid><pubDate>Sun, 03 Sep 2023 12:13:00 GMT</pubDate><content:encoded>&lt;p&gt;Do you ever stop to think about the tiny marvels of technology that make our digital world go &apos;round? We&apos;re diving into the building blocks of the tech universe, unraveling the mysteries behind what enables you to read this very article. We&apos;re talking DNS, ephemeral ports, proxies, tunnels, and more—these are the unsung heroes ensuring data flows seamlessly from one device to another on your network.&lt;/p&gt;
&lt;p&gt;Today, our spotlight is on ephemeral ports. These unsung heroes are the reason you can have multiple online conversations at once. Imagine only being able to chat with one friend at a time—how dull would that be?&lt;/p&gt;
&lt;h2 id=&quot;the-basics-can-you-juggle-two-conversations&quot;&gt;&lt;strong&gt;The Basics: Can You Juggle Two Conversations?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Let&apos;s put you in the spotlight. Have you ever chatted with more than one friend simultaneously? Of course, you have! Picture a scenario where you&apos;re conversing with each of them on different topics. Now, think about how you manage that. Let&apos;s start with a situation where these conversations are happening over text. It&apos;s as simple as opening separate chat windows for each conversation. You effortlessly switch between these windows, seamlessly adding to each discussion. With minimal effort, you maintain the context of each conversation and respond appropriately. The topics can vary widely, but you&apos;ve got it all under control.&lt;/p&gt;
&lt;p&gt;Now, shift gears and imagine you&apos;re doing this in person. Not so challenging, is it?&lt;/p&gt;
&lt;p&gt;Replace those chat windows with the faces of your friends, and you can still smoothly engage in two or more parallel conversations at the same time.&lt;/p&gt;
&lt;p&gt;Let&apos;s dissect what&apos;s happening here: You initiate each conversation, creating a context (even if it&apos;s subconscious) for that specific chat. Without breaking a sweat, you receive responses and place them neatly within the corresponding context. When it&apos;s time to respond, you effortlessly switch back to the relevant context and continue the conversation.&lt;/p&gt;
&lt;h2 id=&quot;ephemeral-ports-where-your-tech-devices-get-playful&quot;&gt;&lt;strong&gt;Ephemeral Ports: Where Your Tech Devices Get Playful&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Have you ever wondered how your tech gadgets, be it your mobile, laptop, or even your trusty smartwatch, manage to mimic human-like communication? Well, behind the scenes, they rely on something called &quot;ephemeral ports&quot; to keep their digital conversations flowing smoothly. But before we dive into this world of tech wizardry, let&apos;s take a quick detour to refresh our memory on what ports are all about.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ports: The Gateways to Digital Conversations&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Think of ports as the doors to a bustling digital building, the faces in a crowd, or even the chat windows in our digital lives. They are the identifiers that machines use to keep track of different ongoing conversations. Every message that goes out or comes in gets neatly sorted into these virtual buckets, and it&apos;s the operating system&apos;s job to manage and process the data within them like a maestro conducting a symphony.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Numbers, Numbers Everywhere&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Now, here&apos;s where it gets interesting. When you type in a web address like &quot;&lt;strong&gt;&lt;a href=&quot;https://mayankraj.com/&quot;&gt;https://mayankraj.com/&lt;/a&gt;&lt;/strong&gt;&quot; (a little shameless plug never hurts, right?), you&apos;re actually making a request to &quot;&lt;strong&gt;&lt;a href=&quot;https://mayankraj.com/&quot;&gt;https://mayankraj.com:443/&lt;/a&gt;&lt;/strong&gt;&quot;. That &quot;:443&quot; at the end? That&apos;s the port number! These numbers range from 0 to 65,353. The first 1023 are reserved for super common TCP/IP applications, aptly named &quot;well-known ports.&quot; There are a few global favorites, like 22 for Secure Socket Shell (SSH), 80 for Hypertext Transfer Protocol (HTTP), 443 for the super-secure Hypertext Transfer Protocol Secure (HTTPS), and many more. Then come the &quot;registered ports,&quot; ranging from 1,024 to 49,151, where applications on your operating system can request to listen in. Ports like 8000 or 8080 are often used for local development. Finally, the last block, from 49,152 to 65,535, houses the dynamic ports, reserved for short-lived, on-the-fly communications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Ephemeral Dance&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In the grand scheme of things, when your favorite app wants to talk to anything else, be it on your local network or across the vast internet, it starts by asking the operating system for a free port. Usually, it goes for one from the dynamic port range, but it can also venture into the reserved territory. The OS kindly grants it a port, let&apos;s say 57000, which means any data received on this port will be sent straight to that application.&lt;/p&gt;
&lt;p&gt;Now, when our app wants to chat with another digital entity (say, a web server), it makes a request on the well-known port of the destination (like 443 for HTTPS). But here&apos;s the fun part – it also includes the port number it received from the OS as its return address, which is 57000 in our case. The destination, our trusty web server, gets the message on port 443, processes it gracefully, constructs a reply, and sends it off to the port number 57000. When the OS on the other end receives this data packet on 57000, it&apos;s like the final act in a fantastic play; it forwards it to the waiting application. This special 57000 port? That, my friends, is what we call an &quot;ephemeral port.&quot;&lt;/p&gt;
&lt;p&gt;So, the next time you see your tech devices communicating seamlessly, remember the ephemeral ports doing a lively dance behind the scenes, ensuring that your digital world stays beautifully connected. It&apos;s all part of the delightful play that is modern technology! 🎉🌐&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Data archival on cloud and how to not do it]]></title><description><![CDATA[Making data-driven decisions is the new norm. It won't come as a surprise that the rate at which we are generating data has gone up. As the…]]></description><link>https://mayankraj.com/blog/data-archival</link><guid isPermaLink="false">https://mayankraj.com/blog/data-archival</guid><pubDate>Fri, 02 Jun 2023 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;Making data-driven decisions is the new norm. It won&apos;t come as a surprise that the rate at which we are generating data has gone up. As the volume of data continues to grow exponentially, businesses are increasingly turning to cloud storage for efficient and scalable archiving solutions. The cloud offers numerous advantages, such as cost-effectiveness, flexibility, and easy accessibility. However, archiving data in the cloud requires careful consideration and planning to ensure data integrity, security, and long-term accessibility. In this blog post, we will explore some common mistakes to avoid when archiving data in the cloud.&lt;/p&gt;
&lt;p&gt;In this blog, we will look at Amazon S3 Glacier which is a popular cloud storage service provided by Amazon Web Services (AWS) that offers durable, secure, and cost-effective storage for long-term data archiving and backup. While S3 Glacier provides numerous benefits, it is important to understand the scenarios in which it may not be the ideal choice. In this blog post, we will delve into the strengths of S3 Glacier and highlight situations where alternative storage solutions might be more appropriate.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;key-areas-to-look-out-for&quot;&gt;Key Areas to Look Out for&lt;/h2&gt;
&lt;p&gt;&lt;ins&gt;&lt;strong&gt;The need for Data Classification&lt;/strong&gt;&lt;/ins&gt; is one of the key mistakes organizations make when archiving data in the cloud. It may seem logical but try to think about the last time your team put in conscious effort in classifying the data properly. Not all data is created equal, and without proper classification, it becomes difficult to prioritize data for archiving, apply appropriate retention policies, and allocate storage resources effectively. Invest time in understanding your data, identifying its value, and categorizing it based on its sensitivity, compliance requirements, and business importance.&lt;/p&gt;
&lt;p&gt;Overlooking the &lt;ins&gt;&lt;strong&gt;backup strategy&lt;/strong&gt;&lt;/ins&gt; often leads to bad experiences down the line. Archiving data does not mean it&apos;s exempt from the risk of loss or corruption. Failing to implement a robust backup and recovery strategy is a grave mistake. Many organizations assume that cloud service providers automatically handle backups, but this is not always the case. Cloud providers may offer infrastructure-level redundancy, but it&apos;s essential to have your data backup strategy in place to protect against accidental deletion, data corruption, or service provider failures. Regularly test your backup and recovery processes to ensure they are effective and reliable.&lt;/p&gt;
&lt;p&gt;When in cloud &lt;ins&gt;&lt;strong&gt;compliance&lt;/strong&gt;&lt;/ins&gt; becomes even more important. Compliance regulations and legal requirements dictate how long certain data must be retained. Ignoring or overlooking these policies when archiving data can lead to serious consequences, such as legal liabilities or financial penalties. Ensure you understand the specific data retention requirements for your industry and region. Implement proper retention policies and procedures, and regularly review and update them to stay compliant with evolving regulations.&lt;/p&gt;
&lt;p&gt;You might also be overlooking data validation and integrity checks. &lt;ins&gt;&lt;strong&gt;Data integrity&lt;/strong&gt;&lt;/ins&gt; is crucial for successful long-term archiving. Neglecting to perform regular data validation and integrity checks can result in silent data corruption that goes undetected until it&apos;s too late. Implement checksums, hash functions, or other integrity validation mechanisms to ensure data remains intact and unaltered during the archiving process. Regularly validate archived data to identify and rectify any integrity issues promptly.&lt;/p&gt;
&lt;p&gt;From Day-0, keep an eye on the &lt;ins&gt;&lt;strong&gt;Long-Term Storage Costs&lt;/strong&gt;&lt;/ins&gt;. While cloud storage offers scalability and flexibility, it&apos;s important to consider the long-term costs associated with archiving data. Cloud storage costs can accumulate over time, especially for large-scale archiving projects. Evaluate different storage options, including lower-cost tiers specifically designed for archiving, and consider a mix of storage solutions to optimize cost-effectiveness. Additionally, periodically review and analyze your archiving needs to identify and remove obsolete or unnecessary data.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;s3-glacier&quot;&gt;S3 Glacier&lt;/h2&gt;
&lt;p&gt;S3 Glacier is a highly advantageous option for archiving data in the cloud due to its cost-effectiveness, durability, availability, security, and integration capabilities. It offers a low-cost storage solution for long-term data retention, making it suitable for infrequently accessed data without immediate retrieval requirements. With a remarkable durability of 99.999999999% (11 nines), data stored in S3 Glacier is spread across multiple facilities to ensure high availability. Security is ensured through server-side encryption, access controls, and seamless integration with AWS Identity and Access Management (IAM). Moreover, S3 Glacier seamlessly integrates with S3 lifecycle policies, enabling automated transitions from hot storage tiers to Glacier based on user-defined rules.&lt;/p&gt;
&lt;h2 id=&quot;limitations-and-when-not-to-use-s3-glacier&quot;&gt;Limitations and When Not to Use S3 Glacier:&lt;/h2&gt;
&lt;p&gt;Despite its advantages, there are certain scenarios where S3 Glacier might not be the optimal storage solution.
If you have data that requires frequent or real-time access, S3 Glacier&apos;s retrieval times (ranging from minutes to hours) may not meet your requirements. In such cases, consider using other storage classes like S3 Standard or S3 Intelligent Tiering. S3 Glacier is optimized for large file sizes. If you predominantly work with small files, the overhead associated with Glacier&apos;s minimum storage duration and retrieval costs might outweigh the benefits. Consider other storage options like S3 One Zone-IA or S3 Standard-IA for smaller files.&lt;/p&gt;
&lt;p&gt;If you need short-term storage for data that will be frequently accessed or modified, S3 Glacier is not the appropriate choice. Instead, opt for storage classes like S3 Standard or S3 Intelligent-Tiering, which offer low-latency and high-performance characteristics. While S3 Glacier provides expedited retrieval options, the associated costs can be higher. If you have strict recovery time objectives (RTOs) and need rapid access to your data, alternatives like S3 Standard orIntelligent Tieringring will better suit your needs.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;In conclusion, if you are looking for a cost-effective and scalable solution to manage the increasing volume of data then archiving data in the cloud could be a great option. But to ensure successful archiving, you should make sure to look into a few aspects. Proper data classification is essential for prioritizing data, applying retention policies, and so is effectively allocating storage resources. Implementing a robust strategy to back up the data, including regular testing, protects against data loss and corruption. Compliance with data retention requirements and regular validation checks ensure data integrity and regulatory adherence cannot be ignored. Considering long-term storage costs and periodically reviewing archiving needs optimizes cost-effectiveness.&lt;/p&gt;
&lt;p&gt;Amazon S3 Glacier is a great option but remember that it is just that - an option. I would use Glacier with my eyes closed for certain use cases but certainly not for all of them. Once you know of the limitations of the tool, you can make better and more informed decisions. For some cases, alternative storage classes like S3 Standard or S3 Intelligent Tiering are more appropriate. Short-term storage needs with frequent access or modification are better served by S3 Standard or S3 Intelligent-Tiering, which offer low-latency and high-performance characteristics. Considering recovery time objectives and the associated costs, organizations can make informed decisions about the most suitable storage option for their specific requirements.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Zero Day Vulnerabilities and You : Introduction]]></title><description><![CDATA[You may not have realised it but there are a handful if not more zero day vurlnerabilities running on the very device you are holding…]]></description><link>https://mayankraj.com/blog/zero-day-vulnerabilities-and-you-introduction</link><guid isPermaLink="false">https://mayankraj.com/blog/zero-day-vulnerabilities-and-you-introduction</guid><pubDate>Wed, 15 Feb 2023 16:53:24 GMT</pubDate><content:encoded>&lt;p&gt;You may not have realised it but there are a handful if not more zero day vurlnerabilities running on the very device you are holding. Weather that is a mobile device, laptop, tablet or even your smart fridge. you can get a grip on their seriousness by knowing that Zero day vurlnerabilties are traded for hundreeds if not millions of dollars among Cyber Criminals and nation state actors.&lt;/p&gt;
&lt;p&gt;To name a few, attacks like Wanacry which affected hundreds of thousands of computers in over 150 countries in May 2017, Stuxne which used a sophisticated piece of malware that targeted Iran&apos;s nuclear program and was discovered in 2010 were done with help of one or more Zero Day Vurnerability at it&apos;s core.&lt;/p&gt;
&lt;p&gt;In this series, I will attempt to tear apart a few such attacks. We will look at how the attack was caries out, what the exploit was, it&apos;s impact and how it was eventually fixed. The language of the articles in the series will ensure that any and everyone will be able to appreciate the execution behind these attacks, in makign them happen and fixing them. This is the most facinating part for me and is what I want to share.&lt;/p&gt;
&lt;br/&gt;
&lt;hr&gt;
&lt;!-- ## But first, What is a Zero Day Vulnerability ? --&gt;</content:encoded></item><item><title><![CDATA[Unique Quirks of DynamoDB: Things you will notice when building at scale]]></title><description><![CDATA[DynamoDB is part of the Database family of AWS. It sits beside the heavyweights like RDS, Redshift, etc. It is a fully managed NoSQL…]]></description><link>https://mayankraj.com/blog/dynamodb-quirks</link><guid isPermaLink="false">https://mayankraj.com/blog/dynamodb-quirks</guid><pubDate>Sat, 28 Jan 2023 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;DynamoDB is part of the Database family of AWS. It sits beside the heavyweights like RDS, Redshift, etc. It is a fully managed NoSQL database. The core value of DynamoDB is that it provides a high-performance storage solution for applications that require low-latency access to large amounts of data. Above all, it is scalable and reliable, even at large data volumes in tunes of a couple of hundred gigabytes.&lt;/p&gt;
&lt;p&gt;Above all, it&apos;s a managed solution. While that is great news to start with but also means that you don&apos;t get to finetune it when the need arises. So you have to make sure that from Day 0, right from the time when you are designing the database patterns, you also design for the unique behavior of DynamoDB.&lt;/p&gt;
&lt;br/&gt;
&lt;hr&gt;
&lt;p&gt;&lt;ins&gt;&lt;strong&gt;Flexible Schema&lt;/strong&gt;&lt;/ins&gt; is the headlining feature of the database, like any other NoSQL database. Flexible schemas sound like great news but when you go from a few records to a billion, it can come back to bite you. The flexible schema allows you to store items with varying attributes within the same table. This provides great flexibility when dealing with evolving data structures. For example, let&apos;s say you have an e-commerce application where different product categories have different attributes. With DynamoDB, you can store products of various categories in a single table, without needing to define a fixed schema upfront.
DynamoDB indexes play a big role in how you make use of this schema. Without an index, it&apos;s as good as you are not using a database but reading the whole data for every query from a large json file.&lt;/p&gt;
&lt;p&gt;&lt;ins&gt;&lt;strong&gt;Secondary Indexes&lt;/strong&gt;&lt;/ins&gt; will quickly become your friends when you scale. They provide flexible querying capabilities. These indexes can be global, spanning the entire table, or local, limited to a specific partition. Secondary indexes allow you to query your data using different attributes, enhancing the flexibility and efficiency of your application.&lt;/p&gt;
&lt;p&gt;Let&apos;s take an example wherein you are building a social media application and need to retrieve posts by both creation date and user ID. By creating a global secondary index on the user ID attribute, you can efficiently query the DynamoDB table to retrieve all posts made by a particular user. Integrating DynamoDB with AWS AppSync and AWS Amplify can further simplify the development process by providing managed GraphQL APIs for your frontend applications, seamlessly integrating with DynamoDB&apos;s secondary indexes.&lt;/p&gt;
&lt;p&gt;&lt;ins&gt;&lt;strong&gt;Transparent Scaling&lt;/strong&gt;&lt;/ins&gt; is a big draw of DynamoDB. It dynamically adjusts its capacity based on the workload. This eliminates the need for manual provisioning and ensures consistent performance as your application&apos;s demand fluctuates. Scaling can be done both vertically (throughput per partition) and horizontally (number of partitions).&lt;/p&gt;
&lt;p&gt;Suppose you have a real-time analytics application that experiences varying traffic patterns throughout the day. By integrating DynamoDB with Amazon CloudWatch and AWS Application Auto Scaling, you can monitor the workload and automatically adjust the provisioned capacity of your DynamoDB tables. This enables your application to handle high-traffic periods without compromising performance or incurring unnecessary costs during low-traffic periods.&lt;/p&gt;
&lt;p&gt;&lt;ins&gt;&lt;strong&gt;Item Time to Live&lt;/strong&gt;&lt;/ins&gt; is a unique but underrated feature, which automatically deletes expired items from a table. This can be useful for managing temporary data or purging stale records, reducing storage costs and query overhead. This allows you to think of DynamoDB for a lot more use cases than just durable long-term data storage.
Consider a mobile gaming application that tracks user session data. By setting a TTL attribute on the session records in DynamoDB, you can ensure that expired sessions are automatically deleted from the table. Furthermore, you can use DynamoDB Streams in conjunction with AWS Lambda to trigger additional actions whenever an item is deleted, such as updating analytics or sending notifications.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;DynamoDB, in general, seems like a one size fits all, but it is far from it. While it offers several unique quirks that make it a powerful choice for scalable, low-latency data storage but it comes at. cost. Its flexible schema, automatic scaling, secondary indexes, and Time to Live feature provide developers with powerful tools to build efficient and dynamic applications. By combining DynamoDB with other AWS services, you can unleash its full potential and create robust, scalable solutions to meet your application&apos;s specific requirements.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Serverless Architecture 101: Key areas to look into from Day 0]]></title><description><![CDATA[You might have heard about Serverless recently. It has become increasingly popular, with more and more businesses adopting it as a way to…]]></description><link>https://mayankraj.com/blog/serverless-architecture-101</link><guid isPermaLink="false">https://mayankraj.com/blog/serverless-architecture-101</guid><pubDate>Mon, 05 Dec 2022 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;You might have heard about Serverless recently. It has become increasingly popular, with more and more businesses adopting it as a way to build and deploy applications. Moreover, developers are finding it easy to build applications with it. Serverless computing allows developers to focus on writing code, without worrying about the underlying infrastructure.&lt;/p&gt;
&lt;p&gt;Over the years, I&apos;ve built many applications with serverless components embedded into them. Today I prefer to use my serverless templates for even my hobby projects let alone API endpoints that process millions of requests. There are a few key areas that I always find myself coming back to. In this post, I&apos;ve collected 10 such areas that you should pay attention to when building serverless applications, especially using AWS services.&lt;/p&gt;
&lt;br/&gt;
&lt;hr&gt;
&lt;p&gt;Let&apos;s start with the three basics - API, Compute, and Storage.&lt;/p&gt;
&lt;h2 id=&quot;1-api-endpoints&quot;&gt;1. API Endpoints&lt;/h2&gt;
&lt;p&gt;API Gateway is a fully managed service that makes it easy for developers to create, publish, maintain, monitor, and secure APIs at any scale. It acts as a “front door” for your serverless application, allowing you to define RESTful APIs and route requests to backend services. However, if GraphQL is more of your thing, look no further than AWS AppSync.&lt;/p&gt;
&lt;p&gt;For Example: If you want to build a serverless application that requires an API, you can use AWS API Gateway to define your RESTful APIs and route requests to your Lambda functions.&lt;/p&gt;
&lt;h2 id=&quot;2-compute&quot;&gt;2. Compute&lt;/h2&gt;
&lt;p&gt;No brownie points for guessing it. Compute is the foundation of serverless architecture. AWS provides two main computing services: AWS Lambda and AWS Fargate. AWS Lambda is a computing service that lets you run code without provisioning or managing servers. Fargate is a serverless compute engine for containers that allows you to run containers without managing servers or clusters.
For Example: If you want to build a serverless application that requires a lot of computing power, you can use AWS Lambda to run your code. If you have your containers with you fargate may be better suited for you.&lt;/p&gt;
&lt;h2 id=&quot;3-storage&quot;&gt;3. Storage&lt;/h2&gt;
&lt;p&gt;Serverless applications require storage for application data, logs, and other assets. AWS provides several storage services, including Amazon S3, Amazon DynamoDB, and Amazon Aurora serverless.
For Example: If you want to build a serverless application that requires a database, you can use Amazon DynamoDB to store your data.&lt;/p&gt;
&lt;br/&gt;
&lt;hr&gt;
&lt;p&gt;Now with basics out of the way, you have your POC in place. We now get into the interesting aspects of making this application ready for the wide and interesting audience of the public internet.&lt;/p&gt;
&lt;h2 id=&quot;5-breaking-down-the-core-logic&quot;&gt;5. Breaking down the core logic:&lt;/h2&gt;
&lt;p&gt;Break down your application into smaller, focused functions to take full advantage of serverless scalability. Avoid creating monolithic functions that handle multiple tasks.
For example, in an e-commerce application, a lambda function for processing orders and another for sending email notifications would be ideal. You can club them with SQS and make the two logical flows async.&lt;/p&gt;
&lt;h2 id=&quot;6-events-and-more-events&quot;&gt;6. Events and more Events:&lt;/h2&gt;
&lt;p&gt;Always think of breaking into smaller modules and syncing them together with events. Make use of the event-driven architecture to trigger serverless functions. Utilize AWS services such as Amazon S3, Amazon DynamoDB, or Amazon Simple Notification Service (SNS) to trigger functions based on specific events.
For instance, when a new image is uploaded to an S3 bucket, you can automatically resize and optimize it using AWS Lambda.&lt;/p&gt;
&lt;h2 id=&quot;7-data-and-state-durability&quot;&gt;7. Data and State Durability:&lt;/h2&gt;
&lt;p&gt;Serverless functions are inherently stateless, which means they do not maintain a session state. Use managed services like Amazon DynamoDB, Amazon Aurora Serverless, or Amazon Simple Queue Service (SQS) to persist and manage application state across invocations.&lt;/p&gt;
&lt;h2 id=&quot;8-balance-scalability-with-concurrency-&quot;&gt;8. Balance Scalability with Concurrency :&lt;/h2&gt;
&lt;p&gt;Design your serverless application to handle concurrent invocations effectively. Configure the maximum concurrency limits for your functions to avoid resource exhaustion. AWS provides services like AWS Auto Scaling and Amazon API Gateway to automatically scale your serverless application based on demand.&lt;/p&gt;
&lt;h2 id=&quot;9-design-for-cold-starts&quot;&gt;9. Design for Cold Starts:&lt;/h2&gt;
&lt;p&gt;Serverless functions may experience latency due to cold starts when invoked infrequently. Employ strategies such as function warmers, provisioned concurrency, or keeping functions warm with periodic invocations. AWS Lambda provides provisioned concurrency to keep functions ready for instant response.&lt;/p&gt;
&lt;h2 id=&quot;10-distributed-tracing-and-monitoring&quot;&gt;10. Distributed Tracing and Monitoring:&lt;/h2&gt;
&lt;p&gt;Ensure visibility into your serverless application by implementing distributed tracing and monitoring. AWS X-Ray allows you to trace requests as they flow across different serverless functions, helping you identify performance bottlenecks and optimize your application.&lt;/p&gt;
&lt;h2 id=&quot;11-security-and-access-control&quot;&gt;11. Security and Access Control:&lt;/h2&gt;
&lt;p&gt;Implement proper security measures to protect your serverless applications. Leverage AWS Identity and Access Management (IAM) for fine-grained access control. Apply security best practices, such as using secure API gateways, encrypting data at rest and in transit, and following least privilege principles.&lt;/p&gt;
&lt;h2 id=&quot;12-error-handling-and-retry-mechanisms&quot;&gt;12. Error Handling and Retry Mechanisms:&lt;/h2&gt;
&lt;p&gt;Design your serverless application to handle errors gracefully. Utilize features like AWS Step Functions for building resilient workflows, or implement retries with exponential backoff to handle transient failures. AWS Simple Notification Service (SNS) and Amazon Simple Queue Service (SQS) can be used for reliable event processing.&lt;/p&gt;
&lt;h2 id=&quot;13-cost-optimization&quot;&gt;13. Cost Optimization:&lt;/h2&gt;
&lt;p&gt;Optimize the cost of running your serverless application. Configure auto-scaling policies based on demand to avoid over-provisioning. Use AWS Cost Explorer and AWS Budgets to monitor and analyze your serverless costs. Additionally, consider using AWS Lambda Layers to share code across functions and reduce duplication.&lt;/p&gt;
&lt;h2 id=&quot;14-integration-with-existing-systems&quot;&gt;14. Integration with Existing Systems:&lt;/h2&gt;
&lt;p&gt;Leverage AWS services like AWS API Gateway and AWS EventBridge to seamlessly integrate your serverless application with existing systems. Use AWS Lambda as an integration layer to connect disparate components of your application architecture.&lt;/p&gt;
&lt;h2 id=&quot;br&quot;&gt;&lt;br/&gt;&lt;/h2&gt;
&lt;p&gt;Serverless is a great tool. But remember that at the end of the day, it&apos;s just another tool. You should not be looking at forcing the tools to do the job. Designing for serverless architecture requires careful consideration of various aspects to ensure scalability, reliability, and cost-effectiveness. By focusing on function granularity, event triggering, state management, scalability, and other key areas, you can harness the full potential of serverless computing. AWS provides a rich set of services that can be used to address these considerations and build highly efficient and scalable serverless applications.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[An Associate Director talks opportunities, flexibility, and work culture at CACTUS Tech]]></title><description><![CDATA[captionless image Reference by Cactus Tech Follow Mayank Raj is the Associate Director, Engineering, at Cactus Labs. He has extensive…]]></description><link>https://mayankraj.com/blog/cactustech-interview</link><guid isPermaLink="false">https://mayankraj.com/blog/cactustech-interview</guid><pubDate>Mon, 16 Aug 2021 19:40:46 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:1350/format:webp/1*mrHPKdRrs3I0jokL2Ui3hQ.jpeg&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/cactus-techblog/the-ownership-on-a-project-is-like-what-youd-see-in-a-startup-an-associate-director-talks-761a97b41c5a&quot;&gt;Reference&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by &lt;a href=&quot;https://medium.com/@cactus_techblog?source=post_page---byline--761a97b41c5a---------------------------------------&quot;&gt;Cactus Tech&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Follow&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://www.linkedin.com/in/mayank9856/&quot;&gt;Mayank Raj&lt;/a&gt; is the Associate Director, Engineering, at &lt;a href=&quot;https://cactusglobal.com/brands/cactus-labs/&quot;&gt;Cactus Labs&lt;/a&gt;. He has extensive experience in building high-performance products powered by AI/ML &amp;#x26; Big Data. He has designed cloud native architecture that can scale and yet remain cost-effective. He has also worked on more experimental applications like AR/VR.&lt;/p&gt;
&lt;h2 id=&quot;you-joined-cactus-when-you-were-still-in-university-how-did-that-work-out&quot;&gt;&lt;strong&gt;You joined CACTUS when you were still in university. How did that work out?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;That’s right. I joined CACTUS right after my second year of engineering. I had a brief stint with an established name in tech in my second-year summer vacations; I was on contract with their HR team and working on a few apps. I started looking for other contractual opportunities.&lt;/p&gt;
&lt;p&gt;I figured out a hack on LinkedIn and applied to some 400 organizations, including CACTUS, within a few minutes. I wasn’t particularly looking at CACTUS. At that point, I hadn’t even heard about it. But I got a call from the recruitment team and an interview was scheduled. Till date, it remains the best 15 lines of JavaScript I have written.&lt;/p&gt;
&lt;p&gt;During my third and fourth year of university, I started working with CACTUS. I joined the two-member UI team where I formed the other half. I used to come into the office right after college every day; I’d spend a couple of hours at the office, some from home, and some from the last bench of my university class. You can say that this was my introduction to “work from anywhere.”&lt;/p&gt;
&lt;h2 id=&quot;you-mentioned-that-you-had-not-even-heard-of-cactus-so-what-drew-you-to-the-organization&quot;&gt;&lt;strong&gt;You mentioned that you had not even heard of CACTUS. So what drew you to the organization?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;CACTUS gave me a lot of flexibility with work hours, and I needed this since I was still in university. They were ok with me visiting the office for 12–14 hours a week and then doing the rest of the work from anywhere.&lt;/p&gt;
&lt;p&gt;Also, the projects that were discussed during the interview seemed interesting; these were things I would have wanted to work on.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;And most importantly, I was not being treated as an intern.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This was one of the most off-putting things I had to face at other places. I was able to work on fairly complex things, but because I was in university, I was being offered a measly stipend and the promise of “exposure.” I didn’t get the concept. CACTUS didn’t treat me like an intern. I was made a good contractual offer that had good projects and strong ownership from the get-go.&lt;/p&gt;
&lt;p&gt;Finally, the fact that I would have a real job even before graduating seemed like a big bragging right. In retrospect, I don’t think I bragged enough about it. :)&lt;/p&gt;
&lt;h2 id=&quot;can-you-describe-your-role-at-cactus-what-is-your-typical-day-like&quot;&gt;&lt;strong&gt;Can you describe your role at CACTUS? What is your typical day like?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;As Associate Director, Engineering, I serve as a conduit between the business and technology sides of the organization. I take the business requirements and translate them for the tech team. I interface with heads of departments, product managers, and vendors.&lt;/p&gt;
&lt;p&gt;I also make sure that our tech roadmap aligns with our business needs and I manage expectations on both sides. My hands-on programming has reduced over the last year, but I do a lot of Proof-Of-Concepts (POCs) and product bootstrapping.&lt;/p&gt;
&lt;p&gt;I am involved at the very early stages of all products. I work on the first couple of POCs as well as the initial MVP. After that, I help set up a team and move on to the next project.&lt;/p&gt;
&lt;p&gt;I also work on the more experimental stuff. I like to work with new technologies and tools and I get to experiment with something completely new. Almost a year ago, I worked on a VR application. Today I work with a team of BigData Engineers.&lt;/p&gt;
&lt;h2 id=&quot;what-is-the-most-exciting-aspect-of-your-role&quot;&gt;&lt;strong&gt;What is the most exciting aspect of your role?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;This is something that we joke about internally — everyone within the team has equal access to almost everything.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;What that means is that I have the flexibility of working in the capacity that I would want to. I don’t have to wait for approvals. I don’t have to wait for someone to give me a green light to start working on something.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If I pick up a problem to work on, I have the freedom to innovate. This has been the case for all the projects I have worked on. This applies to everyone on the team. This essentially means that all of us are constantly looking for ways to improve and at times geek out on the new offerings.&lt;/p&gt;
&lt;h2 id=&quot;you-make-cactus-sound-like-a-startup&quot;&gt;&lt;strong&gt;You make CACTUS sound like a startup.&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;It’s the best of both worlds. The ownership on a project is like what you’d see in a startup. This is especially true in the tech team. But I don’t have the restrictions that a startup would generally have. It’s not that we are directionless or that we don’t have the in-house capabilities to do most of the things.&lt;/p&gt;
&lt;p&gt;One of the things that startups struggle with is that they don’t know whom to reach out to. In our case, if I have, say, a problem with the language automation solutions we are working on, I can leverage the expertise of the massive team of editors and the close to 20 years of experience we have. All of that is one chat away. I can simply drop in a message and get the conversation going.&lt;/p&gt;
&lt;p&gt;The same goes for budgets. If I can justify the usage, I can have a server farm running to meet my needs.&lt;/p&gt;
&lt;p&gt;When &lt;a href=&quot;https://unsilo.ai/&quot;&gt;UNSILO&lt;/a&gt; joined us, they were surprised with the fact that they no longer had to worry about access to subject-matter experts; we have access to Ph.D. holders that we can collaborate with anytime.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:1400/format:webp/1*JSmLjim21EMnzEvnZLhSIQ.png&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;h2 id=&quot;what-has-been-the-one-coolest-biggest-most-interesting-or-most-challenging-problem-that-youve-worked-on&quot;&gt;&lt;strong&gt;What has been the one coolest, biggest, most interesting, or most challenging problem that you’ve worked on?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;It’s difficult to choose. The first major feature launch I was involved in end-to-end was a big success within the tech team and the entire organization. It was very well received by the stakeholders.&lt;/p&gt;
&lt;p&gt;There was also the time when we were migrating to a new workflow management system, and I was put in charge of revamping the client-facing forms that were the revenue funnels. When I was gearing up for my final year exam in university, I was put in charge of revamping some of the most business critical modules. During that whole process, I took a two-week break for the final university exam and came back to work once the exam was over. So while my batchmates were getting ready for 2–3 months of vacation time, I had kickstarted my professional life!&lt;/p&gt;
&lt;p&gt;Then there’s the time we were working on data generation for a tool. We had to process data, which would have taken close to 6 years if done sequentially, and whittle it down to a couple of hours. Around this time, I along with a team member and the CTO were preparing for a trip to the US for that external collaboration. If we didn’t have the data in time, we would essentially not have anything to do while we were in the US. A waste of everyone’s time, money, and effort! That was the most exciting couple of weeks where I was introduced to the problem statement just a month before the plan date. We were actually able to complete the task just a day before our flight.&lt;/p&gt;
&lt;h2 id=&quot;what-would-you-say-to-people-with-your-background-who-probably-find-working-with-a-startup-or-the-big-tech-companies-more-exciting-whats-the-big-draw-about-cactuss-work-culture&quot;&gt;&lt;strong&gt;What would you say to people with your background who probably find working with a startup or the Big Tech companies more exciting? What’s the big draw about CACTUS’s work culture?&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;For me, the most important consideration is the independence that the organization offers. At CACTUS, I am given a problem and the room to ideate and figure things out on my own. I am given access to leaders with experience whom we can consult, but I am responsible for execution. That has led to a lot of learning opportunities and I can proudly claim to have built something.&lt;/p&gt;
&lt;p&gt;CACTUS is at that sweet spot where it’s not too small and facing a funding crunch or too big and saddled with bureaucratic processes. If I think that I can spend some funds on running a pipeline, I don’t have to wait for approvals as long as I can justify it.&lt;/p&gt;
&lt;p&gt;In my free time, I consult with some startups on technical aspects and the conversations are always around saving a few hundred dollars here and there. You are putting in more effort into saving something that doesn’t really make sense. It’s not the kind of environment I would like to be in.&lt;/p&gt;
&lt;p&gt;Many of my friends working with some big names talk about how restrictive the environment is. Too many hoops to jump through before anything can get started. They are told that they are just starting out and hence cannot call the shots. But it’s the opposite for me and my friends at CACTUS.&lt;/p&gt;
&lt;p&gt;A couple of months ago — this is after some 5 years at CACTUS — I had the opportunity of joining a major cloud computing platform. I spoke to a few people who work there to understand the work culture and it didn’t appeal to me.&lt;/p&gt;
&lt;p&gt;Some of my friends have stuck around with some big names even though they don’t enjoy the work. I’d rather work at a place where I can do interesting work.&lt;/p&gt;
&lt;p&gt;A university batchmate who joined CACTUS recently and who has about two years of work experience is the lead engineer of one of our key tech offerings — a product that generates a couple of hundred thousand dollars in revenue. Those are the kinds of opportunities that people can look forward to here.&lt;/p&gt;
&lt;h2 id=&quot;were-hiring&quot;&gt;We’re hiring!&lt;/h2&gt;
&lt;p&gt;CACTUS Tech is looking for passionate innovators, ideators, and problem solvers to join its team and build exciting products and solutions. Learn more: &lt;a href=&quot;https://tech.cactusglobal.io/careers/&quot;&gt;https://tech.cactusglobal.io/careers/&lt;/a&gt;&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Python Building Blocks - Generators]]></title><description><![CDATA[captionless image Reference by Mayank Raj Follow Every programming language has few aspects that always remain in play and go unnoticed…]]></description><link>https://mayankraj.com/blog/python-generators</link><guid isPermaLink="false">https://mayankraj.com/blog/python-generators</guid><pubDate>Thu, 15 Jul 2021 19:40:46 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:3840/format:webp/1*7nMyswfep-1mgq3MCU4R4A.jpeg&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/cactus-techblog/python-building-blocks-generators-f747717a0bf4&quot;&gt;Reference&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by &lt;a href=&quot;https://medium.com/@mayank9856?source=post_page---byline--f747717a0bf4---------------------------------------&quot;&gt;Mayank Raj&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Follow&lt;/p&gt;
&lt;p&gt;Every programming language has few aspects that always remain in play and go unnoticed. They take the front seat only in two scenarios - when you read about it and it clicks, “Hey ! I’ve been using this all along. Looks like there’s a term for it” or have some niche use case that is satisfied by that exact phenomena. There are few such aspects common across all languages like Scopes, Closures etc and few that are language specific. In the Python ecosystem, generators, iterators, comprehensions etc fall in this category. There’s a good chance you have already used them if you have written more than 100 lines in python.&lt;/p&gt;
&lt;p&gt;In this article, a part of a series touching on such details of the languages we will be looking at one such aspect of Python - Generators. Generators are really powerful if used well. Just like any other tool it works the other way around as well, it can be really bad if not used well or used incorrectly. To a great extent, one can also say that &lt;em&gt;generators&lt;/em&gt; are a specialised &lt;em&gt;iterators.&lt;/em&gt; Iterators is another beast that deserves a spot of its’s own. For this post, we will focus on just &lt;em&gt;generators&lt;/em&gt;.&lt;/p&gt;
&lt;h2 id=&quot;before-generators&quot;&gt;&lt;strong&gt;Before Generators…&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Let’s start by looking at a few cases wherein you would want to iterate over a set of objects. You would write something like this…&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;for item in items:  # items is a list of _something_
     process(item)  # We process each of them one by one
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This looks good. You get the job done and move on. Nothing too fancy here.
…or is there? Can it really be that simple?&lt;/p&gt;
&lt;h2 id=&quot;the-bottlenecks&quot;&gt;&lt;strong&gt;The bottlenecks&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Notice how when you are looping over &lt;code&gt;items&lt;/code&gt; you have to have the &lt;code&gt;items&lt;/code&gt; available to the python interpreter. This means the whole object has to be in memory.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;**What if you had to loop over something that is HUGE !
**What if you had to process lines in a 2TB file. You certainly cannot hold the whole of the 2TB in memory to loop over it.&lt;/li&gt;
&lt;li&gt;**What happens when the items are dynamically generated ?
**In such cases external factors might affect the items that have to be processed.
eg, You took a fresh dump of 100K subscribers of your newsletter and started sending out email to them in a &lt;em&gt;loop.&lt;/em&gt; While sending out, you were 2/5th in and someone unsubscribed from the mailing list. Ideally you would not want to take a fresh dump of subscribers after every email sent out nor would you want to send email to this user after they have unsubscribed because of a race condition.&lt;/li&gt;
&lt;li&gt;**What if you had to abort the loop midway ?
**If the items of the loop were processed beforehand to be made available for looping through them, then you have essentially wasted the processing units.&lt;/li&gt;
&lt;li&gt;How do you structure your code such that it’s intuitive to read and is not sprinkled with list aggregations and processes all over the place ?&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;enter---generators&quot;&gt;&lt;strong&gt;Enter - Generators…&lt;/strong&gt;&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;Generator functions allow you to declare a function that behaves like an iterator, i.e. it can be used in a for loop.
(source: &lt;a href=&quot;https://wiki.python.org/moin/Generators&quot;&gt;Python docs&lt;/a&gt;)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Essentially Generators are simple python functions that make it possible to loop through each item of a &lt;em&gt;list&lt;/em&gt; in a truly sequential manner_,_ without the need of having all the items of the list available beforehand. If it took some computation to generate these items then these computations are done during the loop execution rather than before the loop execution.&lt;/p&gt;
&lt;p&gt;Let’s look at a simple example&lt;/p&gt;
&lt;p&gt;Generators are simple case wherein instead of returning the value, we are instead giving a token for it. This token can be exchanged for the actual value and this exchange happens at the run time. Things become more interesting when you go about using them…&lt;/p&gt;
&lt;p&gt;Do you notice how the response is not the value itself but rather a memory reference ?&lt;/p&gt;
&lt;p&gt;At this point, all that the interpreter has done is take the instructions that you have given it and kept it in memory along with every detail that it needs to execute them and give you the correct answer. &lt;strong&gt;It is however holding back on actually executing on this bit of information until you give the green light.&lt;/strong&gt;
Loosely speaking, this can be called lazy execution. Figuratively and literally. It has the same attitude that you had for your university submissions - procrastinate everything to the last moment until it is absolutely necessary to execute on the plan.&lt;/p&gt;
&lt;p&gt;When you finally need the data…&lt;/p&gt;
&lt;h2 id=&quot;usage-patterns-for-generators&quot;&gt;&lt;strong&gt;Usage Patterns for Generators&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;If you had to read a large file and only process a part of it at once, then generators can help you achieve the goal in an easy manner.&lt;/p&gt;
&lt;p&gt;If your operation on the other hand is expensive, takes time or consumes a lot of resources and you have a case wherein if certain conditions are met you can skip performing the operation on the rest of unchecked items, &lt;em&gt;generators&lt;/em&gt; can save you from building your own logic to handle the use case.&lt;/p&gt;
&lt;p&gt;Finally, you may have also noticed that &lt;em&gt;generators&lt;/em&gt; also help in making the code easier to read. Instead of building the logic as a part of your application, it’s a built in functionality at your disposal.&lt;/p&gt;
&lt;h2 id=&quot;antipatterns-for-generators&quot;&gt;&lt;strong&gt;Antipatterns for Generators&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;Just like any tool, &lt;em&gt;generators&lt;/em&gt; do have a flip side as well. You should not throw around &lt;em&gt;generators&lt;/em&gt; for every loop statement in your application but rather assess the situation and then use it wisely.&lt;/p&gt;
&lt;p&gt;Remember that _generator_s are executed when you actually ask for the value. This makes it a good candidate when you need to factor in the most accurate environmental conditions in your execution (eg, the email newsletter that we discussed earlier)&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;However when you need the environments to be locked in and do not do so explicitly, &lt;em&gt;generators&lt;/em&gt; will pick these dynamically. In some cases, this might not be desirable&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For example, in a stock exchange wherein every second the condition changes, the bets placed have to be evaluated with the parameters (or the environment) present when the orders were placed. If you have a backlog of bets to process, you cannot process a bet with new prices.
Let’s see this in action&lt;/p&gt;
&lt;p&gt;Can you guess what the output of the snippet above would be ? Notice how we change &lt;code&gt;theGlobalMultiplier&lt;/code&gt; midway through our processing. Can you do a dryrun to come up with how the &lt;code&gt;multipliedNumber&lt;/code&gt; would turn out ?&lt;/p&gt;
&lt;p&gt;_Followup question: Can you fix the issue ? We expect all the numbers to be multiplied by the &lt;code&gt;_theGlobalMultiplier&lt;/code&gt; at the time of trigger and not execution.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Generators should be treated as a tool. Any good tool has a way to use it optimally, at the right time and under the correct situation. If done so, &lt;em&gt;generators&lt;/em&gt; can not only simplify your logic, but also improve the flow of your code and increase its readability. However you should know its working well enough because if used under the wrong situation, it can go undetected. It becomes very tricky and hard to spot the source of the problem in such cases.&lt;/p&gt;
&lt;p&gt;Next time you are looping over something, ask yourself a few questions…
Is there any space or memory constraints that I should be taking into consideration ? Do I need to process a chunk of data at a time or the whole data ?
Do I know of cases which I can use to exit for the loop early ? How do I make the best use of those to optimise for computation time ?&lt;/p&gt;
&lt;p&gt;Do share the situations in which you have seen a clever use of this gem of python or better yet, used it.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Of Aspirations and Achievements- Front-end Intern to Solutions Architect in 3.5 Years]]></title><description><![CDATA[captionless image Reference by Cactus Tech Look around and the world has a lot of inspiration and motivation to offer. But nothing quite…]]></description><link>https://mayankraj.com/blog/career-update</link><guid isPermaLink="false">https://mayankraj.com/blog/career-update</guid><pubDate>Tue, 14 Jan 2020 19:40:46 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:4800/format:webp/1*9NXY0ZsfWAjhvzF1K8KO_Q.jpeg&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/cactus-techblog/of-aspirations-and-achievements-front-end-intern-to-solutions-architect-in-3-5-years-fbad522c3048&quot;&gt;Reference&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by &lt;a href=&quot;https://medium.com/@cactus_techblog?source=post_page---byline--fbad522c3048---------------------------------------&quot;&gt;Cactus Tech&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Look around and the world has a lot of inspiration and motivation to offer. But nothing quite compares to a story of humble beginnings and great success.&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/@mayank9856&quot;&gt;Mayank’s&lt;/a&gt; is one such story- about faith and risk, about the courage to dive into the unknown and do what’s never been done before, about the right energy focused at the right places, about progress as an outcome of being ever-learning, always-exploring, and growing out-of-the-mould you’ve been set in.&lt;/p&gt;
&lt;h2 id=&quot;the-story-behind-the-story&quot;&gt;The story behind the story&lt;/h2&gt;
&lt;p&gt;Mayank became a part of CACTUS while he was entering the third year of his engineering undergraduate study. He started freelancing a little before that. Initially, it was all about seniors and college committees. &lt;strong&gt;He was “the go-to guy” when it came to developing websites for anything&lt;/strong&gt;, at a time when everyone wanted websites for their next startup or to simply get an online presence. He did an internship with Tata Consultancy Services in the summer break of his first year of college and completed in a week and half, what his manager had planned for him for two months. In the break of his second year, he did a project with Directi. By this time, he had completed multiple projects and built up his portfolio. But working at CACTUS would get him involved in a long-term project and that was a big leap from the unstructured, erratic in-flow of freelancing work. CACTUS seemed like the perfect opportunity to hone his skills and start a career — all at the same time.&lt;/p&gt;
&lt;h2 id=&quot;first-impressions&quot;&gt;First impressions&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;“I was a second-year college student with only some freelance projects and workshop talks to brag about. I didn’t have an under-100 CodeChef rank, some national-level competition medal to show. When I was called for an interview at CACTUS, I did not expect it to be so big. When I got to the office for interview, I was taken aback. The office spanned multiple floors, Cactus had offices in 7 countries, and I saw lot of expats working in the office, and I was interviewed by two people, then team lead and the VP of engineering.” &lt;em&gt;exclaims Mayank, who was surprised at the scale of work done at CACTUS, just like we were amazed at the amount of work he could handle and the pace at which he could deliver results.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;“When I finally joined, I was fortunate enough to have a manager who knew what he was doing and was very good at it. On top of that he was open to my suggestions. Also, I had only heard of terms like ‘no hierarchy at work’, ‘open structure’. I suggested a few things to the then Technical Architect of the team, and in a weeks’ time was working on implementing them.” &lt;em&gt;he recounts.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;so-far-so-great&quot;&gt;So far, so great&lt;/h2&gt;
&lt;p&gt;Mayank joined CACTUS at a stage when the company started using technology not as a support system but to drive transformation and to scale up. This shift, coupled with the fact that even at that time CACTUS was not a small company by any means meant that there was a big list of ‘To-Dos’. This being the perfect environment for him to grow, meant that Mayank would occasionally pick up something exploratory to work on, create a POC, demo it and then start working on it as a full-blown project. Most of the initial big tickets have been this way.&lt;/p&gt;
&lt;p&gt;Given the scale of business, whatever built here is shipped to a huge global audience. Mayank has been able to grow in understanding regarding potential problems that may arise when a product is used by someone in Japan, China, Brazil, USA, etc. and how to handle them, thus, opening up many new avenues for him as a developer.&lt;/p&gt;
&lt;h2 id=&quot;the-developers-ladder&quot;&gt;The developer’s ladder&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;The best thing about CACTUS is that your title doesn’t define your work but rather, it’s the other way around.&lt;/em&gt;&lt;/strong&gt; &lt;em&gt;Mayank recounts,&lt;/em&gt; “Even while in college, I was sleeping during the lectures in the morning and working on improving the then Angular2 Core to accommodate multilingual builds from the same codebase. I had the opportunity to design the architecture and develop the pipeline for a module that processes a great load of events every hour. The project introduced me to AWS, BigData, Analytics and many more things.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;With more responsibility after each project, its quite evident that a company that focuses on merit gained through action, doesn’t make years of experience, the team you’re a part of, or your professional title, a mandate for where you can go.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:1400/format:webp/1*bt3y5cIcdvUUwnlIcYMGSA.jpeg&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“Every other day for me at Cactus is different. Today I get to work on multiple exciting projects in the domain of Artificial Intelligence, NLP (Natural Language Processing), Big Data, cloud architecture and also cutting-edge technologies like Web based Augmented/Virtual Reality. I also have the privilege of managing a small but extremely talented and motivated team, and together we not only bootstrap a product but take it from an idea to a production-ready, scalable and fault-tolerant system. In the process we have learned what it takes to build, scale, and manage large scale applications. I also manage recruitment, take interviews for the team and conduct campus placements. I visited my own university campus in just under two years of graduation, only this time, on the other side of the table to recruit students for my team.&lt;/p&gt;
&lt;p&gt;I’ve also visited NYC, USA, multiple conferences in India and much more. I’ve interacted with core members of the engineering team from AWS, GCP, etc. for various products. In the process I’ve met some brilliant people — both within CACTUS and outside, who are amazing at what they do.”, &lt;em&gt;he remarks.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Some people learn and deliver more in two years than others do in a decade. It’s not the time that matters, it’s the individual, and CACTUS gives you the right opportunities and freedom to turn your aspirations into your achievements.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Demystifying JavaScript Closures]]></title><description><![CDATA[captionless image Reference by Mayank Raj Follow JavaScript as a language is easy to get started with but difficult to master. This is…]]></description><link>https://mayankraj.com/blog/javascript-closures</link><guid isPermaLink="false">https://mayankraj.com/blog/javascript-closures</guid><pubDate>Sat, 30 Nov 2019 19:40:46 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:4000/format:webp/1*3ahzyT4QAHsX_oVoTDv__w.jpeg&quot; alt=&quot;captionless image&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/cactus-techblog/demystifying-javascript-closures-2628c807bf18&quot;&gt;Reference&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by &lt;a href=&quot;https://medium.com/@mayank9856?source=post_page---byline--2628c807bf18---------------------------------------&quot;&gt;Mayank Raj&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Follow&lt;/p&gt;
&lt;p&gt;JavaScript as a language is easy to get started with but difficult to master. This is partly owing to the fact that you can get up and running with a fairly complex application built with JS and not know in detail the workings of it and somewhat because once you have a functional system there is not much motivation to go back and break it down. Hoisting, Lexical Scopes, Closures, &lt;code&gt;this&lt;/code&gt; and even IIFFE are not well understood by developers. Understanding concepts like these not only make the language interesting but also explain a lot of weird - “Oh this works!! I don&apos;t know why or how, but it does!”. Let&apos;s break down Closures today.&lt;/p&gt;
&lt;p&gt;A write up on Closure is not complete without the line, ‘you have already used them before’, so its only fitting that I start with it as well.&lt;/p&gt;
&lt;h2 id=&quot;understanding-lexical-scopes&quot;&gt;Understanding Lexical Scopes&lt;/h2&gt;
&lt;p&gt;A good point to start with would be understanding what scopes actually are. Scope, at the most basic level is exactly what it sounds like — scope of the variables. It’s how the compiler answers the following question -&lt;/p&gt;
&lt;p&gt;&lt;em&gt;I have two variable&lt;/em&gt; &lt;code&gt;_foo_&lt;/code&gt; &lt;em&gt;and&lt;/em&gt; &lt;code&gt;_bar_&lt;/code&gt;&lt;em&gt;, if someone (function, assignment, etc.) asks for it should I give it to the requester? Just to break it to you, yes JavaScript is a compiled language. It all happens at run time and not ahead of time like in the case of C++, Java, etc. so it&apos;s not very evident. Hoisting is one of the proof of this. Anyways, getting back to the topic, a compiler should know where a variable was declared and who has access to that variable. Apart from this being the basic requirement of a program, it is also useful for garbage collecting (once I know no one can access a variable, I can safely delete it). Everything in JS is defined under a scope. That scope can either be of global execution context or that of a function. Every new scope is a bubble inside of its parent scope. Now the crucial part is that each scope has access to everything declared in itself and also in the scope of its parents, ancestors etc. The first place to look for is the local scope, then one step above and so on. Let&apos;s see it in action.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We’ll break down the execution of the above example when we call &lt;code&gt;greet&lt;/code&gt;. There are three scopes in play here - Global execution, that of function &lt;code&gt;greet&lt;/code&gt; and of function &lt;code&gt;printGreet&lt;/code&gt;. The flow of how variables are discovered is as follows. When the execution comes to line 6, there are three variables that are called for here.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;code&gt;heyString&lt;/code&gt;: There is no declaration found in the local scope i.e. that of the function &lt;code&gt;printGreet&lt;/code&gt;. We move to an upper bubble and into the scope of the function &lt;code&gt;greet&lt;/code&gt;. Voila, we got the reference of the variable here.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;name&lt;/code&gt;: Local scope of &lt;code&gt;printGreet&lt;/code&gt; doesn&apos;t have any references of the variable, we go a step above and we get the reference in the function &lt;code&gt;greet&lt;/code&gt; again. This time it was an argument that was passed to the function.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;greetingEmoji&lt;/code&gt;: The story repeats here as well but we don&apos;t find any reference in the scope of the function &lt;code&gt;greet&lt;/code&gt; as well. We know what we have to do and move up in the scope bubble. We are now in the Global Execution scope and here we find a reference of the variable we were looking for. Thus declaring scope is in the hands of the author. It depends on where you declare something i.e. the scope of the variable, the bubble that it is kept in is dependent on the position of its declaration in the program. The nesting of scopes is a byproduct of this functionality. This is what is termed as &lt;code&gt;lexical scope&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;closure&quot;&gt;Closure&lt;/h2&gt;
&lt;p&gt;Now on to the main business. Closure can be summed up as:&lt;/p&gt;
&lt;p&gt;The ability of a function to hold a reference to it’s Lexical scope even when the execution is happening outside it. With a new profound knowledge of Lexical Scopes and the above line, have a look at the below code and try to guess the final output and reason it.&lt;/p&gt;
&lt;p&gt;Let’s make our observations loud and clear:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A function &lt;code&gt;customGreet&lt;/code&gt; accepts an argument which is assigned to a local variable &lt;code&gt;customGreetingPrefix&lt;/code&gt;. It also has a declaration of a function which is eventually returned.&lt;/li&gt;
&lt;li&gt;This declaration is of a function &lt;code&gt;printGreet&lt;/code&gt; which accepts an argument &lt;code&gt;name&lt;/code&gt; and prints a greeting message.&lt;/li&gt;
&lt;li&gt;We execute the function &lt;code&gt;customGreet&lt;/code&gt; twice at line 8 and 9. With that we store the response in the respective variables.&lt;/li&gt;
&lt;li&gt;The response itself is a function, thus these two variables now hold a reference to a function (i.e. &lt;code&gt;printGreet&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;These two functions are then executed. If you notice, the function &lt;code&gt;printGreet&lt;/code&gt; was declared in the lexical scope of &lt;code&gt;customGreet&lt;/code&gt;. But when we finally execute it at line 10 and 11, the lexical scopes have now changed, it is executed in the parent of &lt;code&gt;customGreet&lt;/code&gt;. So going by the theory that access to a variable is only successful if it is present in either the current scope or any of the parent scope, we can say that the function &lt;code&gt;printGreet&lt;/code&gt; would not be able to access &lt;code&gt;customGreetingPrefix&lt;/code&gt;. But to our surprise, it can access it! What you see my friend, is closure in pay here. By definition, a function holds the reference to it&apos;s scope. So in our case, the function &lt;code&gt;printGreet&lt;/code&gt; holds the reference to the lexical scope of &lt;code&gt;customGreet&lt;/code&gt;, which, (you guessed it) brings the reference of &lt;code&gt;customGreetingPrefix&lt;/code&gt; with it (technically also of &lt;code&gt;greetingPrefix&lt;/code&gt;). Thus when we create the scope of &lt;code&gt;customGreet&lt;/code&gt; at line 8 and 9, the two bubbles are preserved by the JS engine as it knows that &lt;code&gt;printGreet&lt;/code&gt; may ask for anything in that scope at any point later. Thus we get the following output.&lt;/li&gt;
&lt;/ol&gt;
&lt;pre&gt;&lt;code&gt;...
greetWithHey(&apos;Tom&apos;);     // Hey Tom.
greetWithHello(&apos;Jerry&apos;); // Hello Jerry.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now that you know the signature of Closure, they are not very hard to find. Every JS module that you use adopts closure in some form or the other. Consider the following example.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function theBestModule(config) {
    var propertyA = config[&apos;propertyA&apos;];
    var propertyB = config[&apos;propertyB&apos;];
    var propertyC = config[&apos;propertyC&apos;];
    def combinePropertyAandB() {
        return (propertyA + propertyB);
    }
    def combinePropertyAandC() {
        return (propertyA + propertyC);
    }
    return {
        &apos;combineAandB&apos;: combinePropertyAandB,
        &apos;combineAandC&apos;: combinePropertyAandC
    }
}
myModulesConfig = { /* ... */ }
myModule = theBestModule(myModulesConfig);
console.log(myModule.combineAandB) // TADA, we found closure
console.log(myModule.combineAandC) // ...yet again. Not that difficult huh.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Virtually all modules in JS accept a set of configuration in some form or the another. Internally, they set these configurations to default and then override it if the user passed something custom. If you notice, these configurations are accessed by methods that the module (a function to be precise) exposes. You don’t have access to these configurations because the lexical scope is not available to you but the functions declared inside the module hold the reference.&lt;/p&gt;
&lt;p&gt;With all the knowledge that you have collected, I’ll leave you with the following program. Try to guess what the output would be at each log statement and reason it.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;for(var i = 0; i&amp;#x3C;5; i++){
    setTimeout(() =&gt; console.log(i), 0); // ?
}
&lt;/code&gt;&lt;/pre&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Lexical scope is not unique to above examples or to some special cases. It comes into play every time you define a function or create a code block. Even a plain and simple function call has closure over its parent scope. It just becomes more evident in some cases and can quickly become confusing to developers.&lt;/p&gt;
&lt;p&gt;You don’t have to know it to use it but the knowledge of the pattern enables you to effectively use it to your benefit. In many cases, it becomes easy to reason why something is happening the way it is. Closure can help you write clean code by best utilizing resources.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Managing Continuous Integration Pipelines with Jenkins]]></title><description><![CDATA[Introduction Jenkins is acclaimed as the "leading open-source automation server," a versatile tool for automating diverse aspects of a…]]></description><link>https://mayankraj.com/blog/cicd-with-jenkins</link><guid isPermaLink="false">https://mayankraj.com/blog/cicd-with-jenkins</guid><pubDate>Tue, 11 Jun 2019 19:40:46 GMT</pubDate><content:encoded>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Jenkins is acclaimed as the &quot;leading open-source automation server,&quot; a versatile tool for automating diverse aspects of a development workflow. It enables the streamlining of processes such as running test cases, deploying the latest builds, or even automating entire Continuous Integration (CI) pipelines. Jenkins can be configured to execute these tasks in various environments, including dedicated containers. Its distributed nature allows for scalable workloads and simultaneous task execution across different environments, such as running test cases for an Angular project in multiple browsers and versions concurrently. With an extensive library of plugins, Jenkins seamlessly integrates and communicates with external services like Git, Slack, and more.&lt;/p&gt;
&lt;p&gt;In this tutorial, you will learn to set up a Jenkins instance. You will also explore how to leverage the modern Blue Ocean plugin interface and integrate with GitHub to automate tests. You can use &lt;a href=&quot;https://github.com/auth0-blog/jenkins-ci-cd&quot;&gt;this GitHub repository&lt;/a&gt; as a reference if needed.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“Learn how to manage Continuous Integration pipelines with Jenkins.”
Tweet This&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;1-installing-jenkins&quot;&gt;1. Installing Jenkins&lt;/h2&gt;
&lt;p&gt;To begin, you&apos;ll need a server to host Jenkins, making it accessible globally to you and any services you integrate with it. At a minimum, this requires a static IP address or a valid DNS name pointing to the server. This domain (or IP) will be used by external services like GitHub to notify Jenkins of new events.&lt;/p&gt;
&lt;p&gt;If you already have a server, you can skip to section 1.2.&lt;/p&gt;
&lt;h3 id=&quot;11-setting-up-a-digitalocean-droplet&quot;&gt;1.1. Setting up a DigitalOcean Droplet&lt;/h3&gt;
&lt;p&gt;DigitalOcean is a cloud provider simplifying the setup of virtual servers, known as droplets. If you don&apos;t have a DigitalOcean account, you can use &lt;a href=&quot;https://m.do.co/c/4541f2180905&quot;&gt;this link&lt;/a&gt; to get $100 in credit over 60 days.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Navigate to the droplets management console and click on &quot;Create Droplet&quot;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Choose an image:&lt;/strong&gt; Select &quot;Ubuntu 18.04&quot; (or a current LTS version).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Choose a plan:&lt;/strong&gt; Start with a basic plan (e.g., $5/month).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Choose a datacenter region:&lt;/strong&gt; Select one closest to you (e.g., Bangalore).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Authentication:&lt;/strong&gt; Add SSH keys if you have them, or you can use password authentication (explained later).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hostname:&lt;/strong&gt; Assign a descriptive hostname (e.g., &lt;code&gt;jenkins-test&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Verify details and click &quot;Create&quot;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Once the droplet is created, you&apos;ll receive its IP address, username, and password via email (if you didn&apos;t use SSH keys). Log into the server using a terminal:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;ssh &amp;#x3C;username&gt;@&amp;#x3C;public_ip_address&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Replace &lt;code&gt;&amp;#x3C;username&gt;&lt;/code&gt; and &lt;code&gt;&amp;#x3C;public_ip_address&gt;&lt;/code&gt; with your droplet&apos;s credentials.&lt;/p&gt;
&lt;p&gt;You might be prompted to add the IP address to your list of known hosts; accept it. Enter your password when prompted. You may be asked to reset the password.&lt;/p&gt;
&lt;p&gt;You now have a virtual server ready for Jenkins installation.&lt;/p&gt;
&lt;h3 id=&quot;12-installing-jenkins&quot;&gt;1.2. Installing Jenkins&lt;/h3&gt;
&lt;p&gt;After accessing your server, update its packages:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;sudo apt update
sudo apt upgrade
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Install the Java 8 runtime environment, which is required by Jenkins:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# Install Java 8
sudo apt install openjdk-8-jdk

# Verify installation
java -version
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;If the output is similar to &lt;code&gt;openjdk version &quot;1.8.0_212&quot;&lt;/code&gt;, proceed to install Jenkins.&lt;/p&gt;
&lt;p&gt;Add the Jenkins repository key to the system:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;wget -q -O - https://pkg.jenkins.io/debian/jenkins.io.key | sudo apt-key add -
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Add the Debian package repository to your system&apos;s sources:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;sudo sh -c &apos;echo deb http://pkg.jenkins.io/debian-stable binary/ &gt; /etc/apt/sources.list.d/jenkins.list&apos;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Update &lt;code&gt;apt&lt;/code&gt; to use the new source:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;sudo apt-get update
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Finally, install Jenkins:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;sudo apt-get install jenkins
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Congratulations! Jenkins is installed. Access your Jenkins setup at &lt;code&gt;http://&amp;#x3C;public_ip_address&gt;:8080&lt;/code&gt;. Jenkins listens on port 8080 by default.&lt;/p&gt;
&lt;h2 id=&quot;2-setting-up-root-user&quot;&gt;2. Setting Up Root User&lt;/h2&gt;
&lt;p&gt;When you first access your Jenkins setup (&lt;code&gt;http://&amp;#x3C;public_ip_address&gt;:8080&lt;/code&gt;), you&apos;ll be asked to &quot;Unlock Jenkins&quot; with a password. This password ensures that the person setting up Jenkins has server access.&lt;/p&gt;
&lt;p&gt;Fetch this password from your server:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;sudo cat /var/lib/jenkins/secrets/initialAdminPassword
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Copy the output and paste it into the Jenkins setup page.&lt;/p&gt;
&lt;p&gt;On the next screen, select &quot;Install suggested plugins.&quot; You&apos;ll see the installation progress.&lt;/p&gt;
&lt;p&gt;After plugin installation, create your first admin user by filling in the required details.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Image: Filling in the details about the first Jenkins user.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Next, provide the domain name Jenkins will use (e.g., &lt;code&gt;jenkins.example.com&lt;/code&gt;). For this tutorial, you can use your server&apos;s public IP address.&lt;/p&gt;
&lt;p&gt;Click &quot;Save and Finish&quot; to complete the setup. You&apos;ll be redirected to your Jenkins dashboard.&lt;/p&gt;
&lt;p&gt;You have successfully installed Jenkins and are ready to connect it with GitHub.&lt;/p&gt;
&lt;h2 id=&quot;3-installing-jenkins-blue-ocean&quot;&gt;3. Installing Jenkins Blue Ocean&lt;/h2&gt;
&lt;p&gt;Jenkins Blue Ocean offers a modern interface for interacting with Jenkins.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;From your Jenkins dashboard, select &quot;Manage Jenkins&quot; from the left navigation bar.&lt;/li&gt;
&lt;li&gt;Select &quot;Manage Plugins.&quot;&lt;/li&gt;
&lt;li&gt;Click on the &quot;Available&quot; tab.&lt;/li&gt;
&lt;li&gt;Search for &quot;Blue Ocean.&quot;&lt;/li&gt;
&lt;li&gt;Check the box next to it and click &quot;Install without restart.&quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;Image: Installing Jenkins Blue Ocean plugin.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Once installation completes, access the new interface at &lt;code&gt;http://&amp;#x3C;public_ip_address&gt;:8080/blue&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&quot;4-setting-up-a-pipeline-and-connecting-to-github&quot;&gt;4. Setting Up a Pipeline and Connecting to GitHub&lt;/h2&gt;
&lt;p&gt;When you open the Blue Ocean URL for the first time, you&apos;ll be prompted to create a new pipeline.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Image: Create your first Blue Ocean pipeline.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;Pipeline&lt;/strong&gt; is a defined set of procedures with tasks like running test cases, verifying code, creating deployment packages, and deploying to servers. Triggers initiate these tasks, and each run is called a &lt;strong&gt;Job&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The goal is to use GitHub actions (like commits) as triggers.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Click &quot;Create a new Pipeline.&quot;&lt;/li&gt;
&lt;li&gt;Select &quot;GitHub.&quot;&lt;/li&gt;
&lt;li&gt;You&apos;ll need to create a GitHub access token. Click the &quot;Create access token here&quot; link.
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Note:&lt;/strong&gt; &quot;Jenkins Integration&quot; (or similar).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Select scopes:&lt;/strong&gt; Leave with default values.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Click &quot;Generate token.&quot; Copy this token (GitHub shows it only once).&lt;/li&gt;
&lt;li&gt;Paste the token into the &quot;Your GitHub access token&quot; field in Blue Ocean and click &quot;Connect.&quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Jenkins will now ask you to select a GitHub organization and repository.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Open GitHub in a new tab and create a new repository named &lt;code&gt;jenkins-test&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Back in Jenkins:
&lt;ul&gt;
&lt;li&gt;Select the organization where the repo resides.&lt;/li&gt;
&lt;li&gt;Select the &lt;code&gt;jenkins-test&lt;/code&gt; repository.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Click &quot;Create Pipeline.&quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;Image: Using Jenkins to create a new pipeline on a GitHub repository.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Your pipeline will be ready in a few seconds, and you&apos;ll be redirected to the pipeline editor.&lt;/p&gt;
&lt;h2 id=&quot;5-adding-stages-to-the-pipeline&quot;&gt;5. Adding Stages to the Pipeline&lt;/h2&gt;
&lt;p&gt;A &lt;strong&gt;Stage&lt;/strong&gt; is a logical block in a pipeline housing various &lt;strong&gt;Steps&lt;/strong&gt; (individual actions).&lt;/p&gt;
&lt;p&gt;For this tutorial, the pipeline will check for the presence of two files: &lt;code&gt;file-1.txt&lt;/code&gt; and &lt;code&gt;file-2.txt&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;First, define the &lt;strong&gt;Agent&lt;/strong&gt; (the environment where the pipeline runs). Jenkins supports distributed architectures with master and agent nodes, or Docker containers. For this tutorial, Jenkins will run the pipeline on the host server itself.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In the pipeline editor, select &quot;any&quot; from the &quot;Agent&quot; dropdown on the right.&lt;/li&gt;
&lt;li&gt;Leave environment variables blank for now.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To add a stage:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Click the plus (+) button next to &quot;Start.&quot;&lt;/li&gt;
&lt;li&gt;Name the stage (e.g., &quot;Check file 1&quot;).&lt;/li&gt;
&lt;li&gt;Click &quot;Add Step.&quot;&lt;/li&gt;
&lt;li&gt;Choose step type: &quot;Shell Script.&quot;&lt;/li&gt;
&lt;li&gt;Enter the command: &lt;code&gt;cat file-1.txt&lt;/code&gt; (this checks if the file exists; if not, the step fails).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Similarly, add a second stage named &quot;Check file 2&quot; with a &quot;Shell Script&quot; step using the command &lt;code&gt;cat file-2.txt&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Image: Adding multiple stages to a Jenkins job.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Save this pipeline to your repository. Pipelines are configurations stored in a &lt;code&gt;Jenkinsfile&lt;/code&gt; at the root of your repository. You can also write this file directly (&lt;a href=&quot;https://jenkins.io/doc/book/pipeline/jenkinsfile/&quot;&gt;learn more here&lt;/a&gt;).&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Click &quot;Save&quot; on the top right.&lt;/li&gt;
&lt;li&gt;Add a commit message (e.g., &quot;configuring Jenkins&quot;).&lt;/li&gt;
&lt;li&gt;Select the &lt;code&gt;master&lt;/code&gt; branch.&lt;/li&gt;
&lt;li&gt;Click &quot;Save and Run.&quot; This commits the &lt;code&gt;Jenkinsfile&lt;/code&gt; to your branch and triggers the build.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You&apos;ve defined pipeline stages and triggered your first build.&lt;/p&gt;
&lt;h2 id=&quot;6-setting-up-github-webhook&quot;&gt;6. Setting Up GitHub Webhook&lt;/h2&gt;
&lt;p&gt;To automate builds upon new commits:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Go to your GitHub repository (&lt;code&gt;jenkins-test&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Click the &quot;Settings&quot; tab.&lt;/li&gt;
&lt;li&gt;In the left navigation bar, select &quot;Webhooks.&quot;&lt;/li&gt;
&lt;li&gt;Click &quot;Add Webhook.&quot;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Payload URL:&lt;/strong&gt; &lt;code&gt;http://&amp;#x3C;public_ip_address&gt;:8080/github-webhook/&lt;/code&gt; (replace &lt;code&gt;&amp;#x3C;public_ip_address&gt;&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Leave other settings as default and click &quot;Add Webhook.&quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;Image: Creating a GitHub webhook to notify Jenkins about new commits.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;GitHub will now notify Jenkins of any new commits.&lt;/p&gt;
&lt;h2 id=&quot;7-triggering-the-first-automated-build&quot;&gt;7. Triggering the First Automated Build&lt;/h2&gt;
&lt;p&gt;Test the integration by making a commit. You can either:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Clone your repository locally, add files, commit, and push.&lt;/li&gt;
&lt;li&gt;Use GitHub&apos;s website to create the files.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Add two files to the project root: &lt;code&gt;file-1.txt&lt;/code&gt; and &lt;code&gt;file-2.txt&lt;/code&gt; (content doesn&apos;t matter for this test).&lt;/p&gt;
&lt;p&gt;If working locally, commit and push the new files. GitHub will notify Jenkins, triggering an automatic build (which should succeed if both files are present).&lt;/p&gt;
&lt;p&gt;If using GitHub&apos;s website, adding the first file will trigger a build (which will likely fail if the second file check is active). Adding the second file will trigger another build, which should then succeed.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Image: Jenkins showing the list of builds that were automatically triggered after GitHub&apos;s webhooks notifications.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;New commits to this repository will now automatically trigger builds, and results will be updated on GitHub.&lt;/p&gt;
&lt;h2 id=&quot;8-recording-test-results-and-artifacts&quot;&gt;8. Recording Test Results and Artifacts&lt;/h2&gt;
&lt;p&gt;Each Jenkins build typically starts with a shallow clone of the branch in a new, isolated folder. After the build, Jenkins performs a cleanup, deleting this folder. To preserve test results or build outputs, use &lt;strong&gt;build artifacts&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Build artifacts are stored in &lt;code&gt;/var/lib/jenkins/workspace/&lt;/code&gt; on your Jenkins server if you&apos;re curious.&lt;/p&gt;
&lt;p&gt;This section transforms your repository into a simple npm project to demonstrate recording test results. (npm/Node.js knowledge isn&apos;t strictly required to follow, but you need them installed locally for these steps.)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Jenkins understands JUnit XML format for test results, a standard supported by most test runners. This tutorial uses MochaJS.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Ensure npm and Node.js are installed locally.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Clone your &lt;code&gt;jenkins-test&lt;/code&gt; repository.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Initialize an npm project:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# from the repository project root
npm init -y
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Install dependencies:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# install the test case runner
npm install --save-dev mocha

# install the JUnit XML reporter
npm install --save-dev mocha-jenkins-reporter
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Create &lt;code&gt;test.js&lt;/code&gt; with the following content:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;const assert = require(&apos;assert&apos;);

describe(&apos;Addition&apos;, function() {
  it(&apos;2 + 2 should be 4&apos;, function() {
    assert.equal((2 + 2), 4);
  });

  it(&apos;1 + 5 should be 6&apos;, function() {
    assert.equal((1 + 5), 6);
  });
});

describe(&apos;Type Comparison&apos;, function() {
  it(&apos;\&apos;5\&apos; == 5 should be true&apos;, function() {
    assert.equal((&apos;5&apos; == 5), true);
  });

  it(&apos;\&apos;5\&apos; === 5 should be false&apos;, function() {
    assert.equal((&apos;5&apos; === 5), false);
  });
});
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Run tests locally:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;npx mocha
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Generate reports in JUnit XML format:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# define where you want the test results
export JUNIT_REPORT_PATH=./test-results.xml

# run mocha and tell it to use the JUnit reporter
npx mocha --reporter mocha-jenkins-reporter
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This saves results to &lt;code&gt;test-results.xml&lt;/code&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Update your GitHub repository:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# make git ignore the node_modules directory
echo &quot;node_modules/&quot; &gt; .gitignore

# stage your new code
git add .

# commit it to git
git commit -m &quot;Added Test Cases&quot;

# push it to GitHub
git push origin master
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, configure Jenkins to run these tests and find the XML report.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Go to the Blue Ocean pipeline editor (e.g., &lt;code&gt;http://&amp;#x3C;public_ip_address&gt;:8080/blue/organizations/jenkins/pipeline-editor/jenkins-ci-cd/master/&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Add a new stage after &quot;Check file 2&quot; called &quot;Install dependencies&quot;.
&lt;ul&gt;
&lt;li&gt;Add a &quot;Shell Script&quot; step: &lt;code&gt;npm install -d&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Add another stage after &quot;Install dependencies&quot; called &quot;Run test cases&quot;.
&lt;ul&gt;
&lt;li&gt;Add a &quot;Shell Script&quot; step:&lt;/li&gt;
&lt;/ul&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;# define where you want the test results
export JUNIT_REPORT_PATH=./test-results.xml

# run mocha and tell it to use the JUnit reporter
npx mocha --reporter mocha-jenkins-reporter
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;Click &quot;Save&quot; on the top right.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Install npm and Node.js on your Jenkins server (as global dependencies):&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-bash&quot;&gt;ssh &amp;#x3C;username&gt;@&amp;#x3C;public_ip_address&gt;
# Then, on the server:
sudo apt install nodejs npm
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To capture JUnit-formatted test results, you need to update the &lt;code&gt;Jenkinsfile&lt;/code&gt; directly, as the visual editor might not support the &lt;code&gt;post&lt;/code&gt; stage for this.
Open &lt;code&gt;Jenkinsfile&lt;/code&gt; (at the project root) and update it:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-groovy&quot;&gt;pipeline {
  agent any
  stages {
    // ... all existing stages remain untouched ...
    stage(&apos;Install dependencies&apos;) {
        steps {
            sh &apos;npm install -d&apos;
        }
    }
    stage(&apos;Run test cases&apos;) {
        steps {
            sh &apos;&apos;&apos;
            export JUNIT_REPORT_PATH=./test-results.xml
            npx mocha --reporter mocha-jenkins-reporter
            &apos;&apos;&apos;
        }
    }
  }
  post {
    always {
        junit &apos;test-results.xml&apos;
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Commit and push this change. You should now find test results on the &quot;Tests&quot; tab of your Jenkins pipeline.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Image: Checking tests results on a Jenkins pipeline.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;You&apos;ve learned to record test results and track their status across builds.&lt;/p&gt;
&lt;h2 id=&quot;9-exploring-plugins&quot;&gt;9. Exploring Plugins&lt;/h2&gt;
&lt;p&gt;Plugins are the backbone of Jenkins&apos; power and flexibility. Explore available plugins &lt;a href=&quot;https://plugins.jenkins.io/&quot;&gt;here&lt;/a&gt;.
Examples:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Slack Notification:&lt;/strong&gt; Notify your Slack channel about build statuses.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kubernetes Plugin:&lt;/strong&gt; Integrate with Kubernetes.&lt;/li&gt;
&lt;li&gt;Cloud container service plugins like &lt;strong&gt;Amazon Elastic Container Service&lt;/strong&gt; and &lt;strong&gt;Azure Container Service&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;“I just built a Continuous Integration pipeline with Jenkins. So cool!”
Tweet This&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;aside-syncing-authentication-in-ci-using-the-auth0-deploy-cli-tool&quot;&gt;Aside: Syncing Authentication in CI using the Auth0 Deploy CLI Tool&lt;/h2&gt;
&lt;p&gt;Managing authentication configuration across different environments (e.g., testing, production) can be tricky. Auth0 can help manage this part of your Jenkins pipeline using the &lt;a href=&quot;https://auth0.com/docs/deploy/deploy-cli-tool&quot;&gt;Auth0 Deploy CLI tool&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The Deploy CLI tool allows you to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Import and export Auth0 tenant configuration objects (tenant settings, rules, connections).&lt;/li&gt;
&lt;li&gt;Export data to a predefined directory structure or a YAML configuration file.&lt;/li&gt;
&lt;li&gt;Call the tool programmatically.&lt;/li&gt;
&lt;li&gt;Replace environment variables.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To try this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://auth0.com/signup&quot;&gt;Sign up for a free Auth0 account&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Check the documentation on &lt;a href=&quot;https://auth0.com/docs/deploy/deploy-cli-tool/install-configure-the-deploy-cli&quot;&gt;how to install the CLI tool&lt;/a&gt; and &lt;a href=&quot;https://auth0.com/docs/deploy/deploy-cli-tool/incorporate-the-deploy-cli-into-the-build-environment&quot;&gt;incorporate it into your build environment&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; This tool can be destructive to your Auth0 tenant. Please read the documentation and test on a development tenant before using it in production.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Congratulations! You now have a Jenkins setup capable of automatically testing your codebase with each push. You can track build history, identify when test cases fail, and perform integration testing after pull request merges.&lt;/p&gt;
&lt;p&gt;What are your thoughts? Will you use Jenkins in production? Let us know in the comment box below.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[How To Build a Search Bar with RxJS]]></title><description><![CDATA[Introduction Reactive Programming is a paradigm concerned with asynchronous data streams, in which the programming model considers…]]></description><link>https://mayankraj.com/blog/searchbar-with-rxjs</link><guid isPermaLink="false">https://mayankraj.com/blog/searchbar-with-rxjs</guid><pubDate>Wed, 17 Apr 2019 19:40:46 GMT</pubDate><content:encoded>&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Reactive Programming is a paradigm concerned with asynchronous data streams, in which the programming model considers everything to be a stream of data spread over time. This includes keystrokes, HTTP requests, files to be printed, and even elements of an array, which can be considered to be timed over very small intervals. This makes it a perfect fit for JavaScript as asynchronous data is common in the language.&lt;/p&gt;
&lt;p&gt;RxJS is a popular library for reactive programming in JavaScript. ReactiveX, the umbrella under which RxJS lies, has its extensions in many other languages like Java, Python, C++, Swift, and Dart. RxJS is also widely used by libraries like Angular and React.&lt;/p&gt;
&lt;p&gt;RxJS’s implementation is based on chained functions that are aware and capable of handling data over a range of time. This means that one could implement virtually every aspect of RxJS with nothing more than functions that receive a list of arguments and callbacks, and then execute them when signaled to do so. The community around RxJS has done this heavy lifting, and the result is an API that you can directly use in any application to write clean and maintainable code.&lt;/p&gt;
&lt;p&gt;In this tutorial, you will use RxJS to build a feature-rich search bar that returns real-time results to users. You will also use HTML and CSS to format the search bar.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;(Image from the original tutorial: Demonstration of Search Bar)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Something as common and seemingly simple as a search bar needs to have various checks in place. This tutorial will show you how RxJS can turn a fairly complex set of requirements into code that is manageable and easy to understand.&lt;/p&gt;
&lt;h2 id=&quot;prerequisites&quot;&gt;Prerequisites&lt;/h2&gt;
&lt;p&gt;Before you begin this tutorial you’ll need the following:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A text editor that supports JavaScript syntax highlighting, such as Atom, Visual Studio Code, or Sublime Text. These editors are available on Windows, macOS, and Linux.&lt;/li&gt;
&lt;li&gt;Familiarity with using HTML and JavaScript together. Learn more in &lt;a href=&quot;https://www.digitalocean.com/community/tutorials/how-to-add-javascript-to-html&quot;&gt;How To Add JavaScript to HTML&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Familiarity with the JSON data format, which you can learn more about in &lt;a href=&quot;https://www.digitalocean.com/community/tutorials/how-to-work-with-json-in-javascript&quot;&gt;How to Work with JSON in JavaScript&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The full code for the tutorial is available on &lt;a href=&quot;https://github.com/do-community/rxjs-search-bar&quot;&gt;Github&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;step-1--creating-and-styling-your-search-bar&quot;&gt;Step 1 — Creating and Styling Your Search Bar&lt;/h2&gt;
&lt;p&gt;In this step, you will create and style the search bar with HTML and CSS. The code will use a few common elements from Bootstrap to speed up the process of structuring and styling the page so you can focus on adding custom elements. Bootstrap is a CSS framework that contains templates for common elements like typography, forms, buttons, navigation, grids, and other interface components. Your application will also use Animate.css to add animation to the search bar.&lt;/p&gt;
&lt;p&gt;First, create a file named &lt;code&gt;search-bar.html&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Next, create the basic structure for your application. Add the following HTML to the new file:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!DOCTYPE html&gt;
&amp;#x3C;html&gt;
  &amp;#x3C;head&gt;
    &amp;#x3C;title&gt;RxJS Tutorial&amp;#x3C;/title&gt;
    &amp;#x3C;!-- Load CSS --&gt;

    &amp;#x3C;!-- Load Rubik font --&gt;

    &amp;#x3C;!-- Add Custom inline CSS --&gt;

  &amp;#x3C;/head&gt;
  &amp;#x3C;body&gt;
      &amp;#x3C;!-- Content --&gt;

      &amp;#x3C;!-- Page Header and Search Bar --&gt;

      &amp;#x3C;!-- Results --&gt;

      &amp;#x3C;!-- Load External RxJS --&gt;

      &amp;#x3C;!-- Add custom inline JavaScript --&gt;
      &amp;#x3C;script&gt;

      &amp;#x3C;/script&gt;
  &amp;#x3C;/body&gt;
&amp;#x3C;/html&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Load the CSS for Bootstrap and Animate.css. Add the following code under the &lt;code&gt;&amp;#x3C;!-- Load CSS --&gt;&lt;/code&gt; comment:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Load CSS --&gt;
&amp;#x3C;link rel=&quot;stylesheet&quot; href=&quot;https://stackpath.bootstrapcdn.com/bootstrap/4.2.1/css/bootstrap.min.css&quot; integrity=&quot;sha384-GJzZqFGwb1QTTN6wy59ffF1BuGJpLSa9DkKMp0DgiMDm4iYMj70gZWKYbI706tWS&quot; crossorigin=&quot;anonymous&quot;&gt;
&amp;#x3C;link rel=&quot;stylesheet&quot; href=&quot;https://cdnjs.cloudflare.com/ajax/libs/animate.css/3.7.0/animate.min.css&quot; /&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This tutorial will use a custom font called Rubik from the Google Fonts library. Load the font by adding the following code under the &lt;code&gt;&amp;#x3C;!-- Load Rubik font --&gt;&lt;/code&gt; comment:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Load Rubik font --&gt;
&amp;#x3C;link href=&quot;https://fonts.googleapis.com/css?family=Rubik&quot; rel=&quot;stylesheet&quot;&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next, add the custom CSS to the page under the &lt;code&gt;&amp;#x3C;!-- Add Custom inline CSS --&gt;&lt;/code&gt; comment. This will style the headings, search bar, and results.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Add Custom inline CSS --&gt;
&amp;#x3C;style&gt;
  body {
    background-color: #f5f5f5;
    font-family: &quot;Rubik&quot;, sans-serif;
  }
  
  .search-container {
    margin-top: 50px;
  }
  .search-container .search-heading {
    display: block;
    margin-bottom: 50px;
  }
  .search-container input,
  .search-container input:focus {
    padding: 16px 16px 16px;
    border: none;
    background: rgb(255, 255, 255);
    box-shadow: 0 2px 4px 0 rgba(0, 0, 0, 0.2), 0 25px 50px 0 rgba(0, 0, 0, 0.1) !important;
  }

  .results-container {
    margin-top: 50px;
  }
  .results-container .list-group .list-group-item {
    background-color: transparent;
    border-top: none !important;
    border-bottom: 1px solid rgba(236, 229, 229, 0.64);
  }

  .float-bottom-right {
    position: fixed;
    bottom: 20px;
    left: 20px;
    font-size: 20px;
    font-weight: 700;
    z-index: 1000;
  }
  .float-bottom-right .info-container .card {
    display: none;
  }
  .float-bottom-right .info-container:hover .card,
  .float-bottom-right .info-container .card:hover {
    display: block;
  }
&amp;#x3C;/style&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Now that the styles are in place, add the HTML for the header and input bar under the &lt;code&gt;&amp;#x3C;!-- Page Header and Search Bar --&gt;&lt;/code&gt; comment:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Content --&gt;
&amp;#x3C;!-- Page Header and Search Bar --&gt;
&amp;#x3C;div class=&quot;container search-container&quot;&gt;
    &amp;#x3C;div class=&quot;row justify-content-center&quot;&gt;
        &amp;#x3C;div class=&quot;col-md-auto&quot;&gt;
        &amp;#x3C;div class=&quot;search-heading&quot;&gt;
            &amp;#x3C;h2&gt;Search for Materials Published by Author Name&amp;#x3C;/h2&gt;
            &amp;#x3C;p class=&quot;text-right&quot;&gt;powered by &amp;#x3C;a href=&quot;https://www.crossref.org/&quot;&gt;Crossref&amp;#x3C;/a&gt;&amp;#x3C;/p&gt;
        &amp;#x3C;/div&gt;
        &amp;#x3C;/div&gt;
    &amp;#x3C;/div&gt;
    &amp;#x3C;div class=&quot;row justify-content-center&quot;&gt;
        &amp;#x3C;div class=&quot;col-sm-8&quot;&gt;
        &amp;#x3C;div class=&quot;input-group input-group-md&quot;&gt;
            &amp;#x3C;input id=&quot;search-input&quot; type=&quot;text&quot; class=&quot;form-control&quot; placeholder=&quot;eg. Richard&quot; aria-label=&quot;eg. Richard&quot; autofocus&gt;
        &amp;#x3C;/div&gt;
        &amp;#x3C;/div&gt;
    &amp;#x3C;/div&gt;
&amp;#x3C;/div&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This uses Bootstrap&apos;s grid system. The search bar has an &lt;code&gt;id&lt;/code&gt; of &lt;code&gt;search-input&lt;/code&gt;, which will be used to bind a listener.&lt;/p&gt;
&lt;p&gt;Next, create a location to display search results. Under the &lt;code&gt;&amp;#x3C;!-- Results --&gt;&lt;/code&gt; comment, add the following:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Results --&gt;
&amp;#x3C;div class=&quot;container results-container&quot;&gt;
    &amp;#x3C;div class=&quot;row justify-content-center&quot;&gt;
        &amp;#x3C;div class=&quot;col-sm-8&quot;&gt;
        &amp;#x3C;ul id=&quot;response-list&quot; class=&quot;list-group list-group-flush&quot;&gt;&amp;#x3C;/ul&gt;
        &amp;#x3C;/div&gt;
    &amp;#x3C;/div&gt;
&amp;#x3C;/div&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;ul&lt;/code&gt; element with the &lt;code&gt;id&lt;/code&gt; &lt;code&gt;response-list&lt;/code&gt; will hold the search results.&lt;/p&gt;
&lt;p&gt;At this point, the &lt;code&gt;search-bar.html&lt;/code&gt; file has its basic structure and styling. In the next step, you will write the JavaScript function to handle search terms and results.&lt;/p&gt;
&lt;h2 id=&quot;step-2--writing-the-javascript&quot;&gt;Step 2 — Writing the JavaScript&lt;/h2&gt;
&lt;p&gt;With the HTML structure formatted, you can write the JavaScript code. This code will serve as the foundation for the RxJS implementation.&lt;/p&gt;
&lt;p&gt;Load the RxJS library by adding the following under the &lt;code&gt;&amp;#x3C;!-- Load RxJS --&gt;&lt;/code&gt; comment in &lt;code&gt;search-bar.html&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;!-- Load RxJS --&gt;
&amp;#x3C;script src=&quot;https://unpkg.com/@reactivex/rxjs@5.0.3/dist/global/Rx.js&quot;&gt;&amp;#x3C;/script&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Inside the &lt;code&gt;&amp;#x3C;script&gt;&lt;/code&gt; tag (under the &lt;code&gt;&amp;#x3C;!-- Add custom inline JavaScript --&gt;&lt;/code&gt; comment), store a reference to the HTML element where results will be displayed:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// Add custom inline JavaScript
const output = document.getElementById(&quot;response-list&quot;);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next, add a function to convert the JSON response from the API into HTML elements. This function will clear previous results and set a delay for search result animation. Add the &lt;code&gt;showResults&lt;/code&gt; function:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;const output = document.getElementById(&quot;response-list&quot;);

function showResults(resp) {
    var items = resp[&apos;message&apos;][&apos;items&apos;];
    output.innerHTML = &quot;&quot;;
    let animationDelay = 0; // Corrected: Use let or var for declaration
    if (items.length == 0) {
        output.innerHTML = &quot;&amp;#x3C;li&gt;Could not find any :(&amp;#x3C;/li&gt;&quot;; // Corrected: Wrap message in &amp;#x3C;li&gt; for consistency
    } else {
        items.forEach(item =&gt; {
        let resultItem = `
        &amp;#x3C;div class=&quot;list-group-item animated fadeInUp&quot; style=&quot;animation-delay: ${animationDelay}s;&quot;&gt;
            &amp;#x3C;div class=&quot;d-flex w-100 justify-content-between&quot;&gt;
            &amp;#x3C;h5 class=&quot;mb-1&quot;&gt;${(item[&apos;title&apos;] &amp;#x26;&amp;#x26; item[&apos;title&apos;]) || &quot;&amp;#x26;lt;Title not available&amp;#x26;gt;&quot;}&amp;#x3C;/h5&gt;
            &amp;#x3C;/div&gt;
            &amp;#x3C;p class=&quot;mb-1&quot;&gt;${(item[&apos;container-title&apos;] &amp;#x26;&amp;#x26; item[&apos;container-title&apos;]) || &quot;&quot;}&amp;#x3C;/p&gt;
            &amp;#x3C;small class=&quot;text-muted&quot;&gt;&amp;#x3C;a href=&quot;${item[&apos;URL&apos;]}&quot; target=&quot;_blank&quot;&gt;${item[&apos;URL&apos;]}&amp;#x3C;/a&gt;&amp;#x3C;/small&gt;
            &amp;#x3C;div&gt; 
            &amp;#x3C;p class=&quot;badge badge-primary badge-pill&quot;&gt;${item[&apos;publisher&apos;] || &apos;&apos;}&amp;#x3C;/p&gt;
            &amp;#x3C;p class=&quot;badge badge-primary badge-pill&quot;&gt;${item[&apos;type&apos;] || &apos;&apos;}&amp;#x3C;/p&gt; 
            &amp;#x3C;/div&gt;
        &amp;#x3C;/div&gt;
        `;
        output.insertAdjacentHTML(&quot;beforeend&quot;, resultItem);
        animationDelay += 0.1;                        
        });
    }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The &lt;code&gt;if&lt;/code&gt; block checks for search results. If results are found, the &lt;code&gt;forEach&lt;/code&gt; loop displays them with an animation.&lt;/p&gt;
&lt;p&gt;You&apos;ve now laid the groundwork for RxJS by creating a function that accepts results and renders them on the page.&lt;/p&gt;
&lt;h2 id=&quot;step-3--setting-up-a-listener&quot;&gt;Step 3 — Setting Up a Listener&lt;/h2&gt;
&lt;p&gt;RxJS deals with data streams. In this project, the stream is the series of characters a user types into the search bar. You will add a listener to the input element.&lt;/p&gt;
&lt;p&gt;Recall the &lt;code&gt;search-input&lt;/code&gt; identifier used for the input field:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-html&quot;&gt;&amp;#x3C;input id=&quot;search-input&quot; type=&quot;text&quot; class=&quot;form-control&quot; placeholder=&quot;eg. Richard&quot; aria-label=&quot;eg. Richard&quot; autofocus&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Create a variable to hold a reference to the &lt;code&gt;search-input&lt;/code&gt; element. This will be the Observable for input events. Add this line within your &lt;code&gt;&amp;#x3C;script&gt;&lt;/code&gt; tags, after the &lt;code&gt;showResults&lt;/code&gt; function:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ... (output variable and showResults function)

let searchInput = document.getElementById(&quot;search-input&quot;);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Use the &lt;code&gt;fromEvent&lt;/code&gt; operator from RxJS to listen for &lt;code&gt;input&lt;/code&gt; events on this DOM element. Add the following line:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ... (output variable and showResults function)

let searchInput = document.getElementById(&quot;search-input&quot;);
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
// ... (more operators will be chained here)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The listener is now set up to notify your code of updates to the input element.&lt;/p&gt;
&lt;h2 id=&quot;step-4--adding-operators&quot;&gt;Step 4 — Adding Operators&lt;/h2&gt;
&lt;p&gt;Operators are pure functions that perform operations on data streams. You&apos;ll use operators for tasks like buffering input, making HTTP requests, and filtering results.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Pluck Target Value:&lt;/strong&gt;
The DOM &lt;code&gt;input&lt;/code&gt; event contains various details. We are interested in the value typed into the target element. Use the &lt;code&gt;pluck&lt;/code&gt; operator:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;let searchInput = document.getElementById(&quot;search-input&quot;);
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Filter by Search Term Length:&lt;/strong&gt;
Set a minimum search term length (e.g., three characters) because shorter terms might not yield relevant results or the user might still be typing. Use the &lt;code&gt;filter&lt;/code&gt; operator:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ...
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
    .filter(searchTerm =&gt; searchTerm.length &gt; 2)
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Debounce Input:&lt;/strong&gt;
To ease the load on the API server, ensure requests are sent only at intervals (e.g., 500ms). Use the &lt;code&gt;debounceTime&lt;/code&gt; operator:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ...
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
    .filter(searchTerm =&gt; searchTerm.length &gt; 2)
    .debounceTime(500)
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Distinct Until Changed (Input):&lt;/strong&gt;
Ignore the search term if it hasn&apos;t changed since the last API call. This further optimizes API calls. Use the &lt;code&gt;distinctUntilChanged&lt;/code&gt; operator:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ...
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
    .filter(searchTerm =&gt; searchTerm.length &gt; 2)
    .debounceTime(500)
    .distinctUntilChanged()
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;SwitchMap for API Calls:&lt;/strong&gt;
Query the API with the search term using RxJS&apos;s AJAX implementation. &lt;code&gt;switchMap&lt;/code&gt; is used to chain AJAX calls. It cancels previous pending requests if a new search term comes in. The &lt;code&gt;map&lt;/code&gt; operator within &lt;code&gt;switchMap&lt;/code&gt; structures the API response. The example uses the Crossref API.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ...
Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
    .filter(searchTerm =&gt; searchTerm.length &gt; 2)
    .debounceTime(500)
    .distinctUntilChanged()
    .switchMap(searchKey =&gt; Rx.Observable.ajax(`https://api.crossref.org/works?rows=50&amp;#x26;query.author=${searchKey}`)
        .map(resp =&gt; ({
            &quot;status&quot; : resp[&quot;status&quot;] == 200,
            &quot;details&quot; : resp[&quot;status&quot;] == 200 ? resp[&quot;response&quot;] : [],
            &quot;result_hash&quot;: Date.now() // A simple way to check if response content changed
        }))
    )
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The response is broken into:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;status&lt;/code&gt;: HTTP status (true if 200).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;details&lt;/code&gt;: The actual response data.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;result_hash&lt;/code&gt;: A timestamp to help detect if results have changed.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Filter Unsuccessful API Responses:&lt;/strong&gt;
Use the &lt;code&gt;filter&lt;/code&gt; operator to only accept successful (status 200) API responses:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ... (previous operators)
    .switchMap(searchKey =&gt; Rx.Observable.ajax(`https://api.crossref.org/works?rows=50&amp;#x26;query.author=${searchKey}`)
        .map(resp =&gt; ({
            &quot;status&quot; : resp[&quot;status&quot;] == 200,
            &quot;details&quot; : resp[&quot;status&quot;] == 200 ? resp[&quot;response&quot;] : [],
            &quot;result_hash&quot;: Date.now()
        }))
    )
    .filter(resp =&gt; resp.status !== false) // or resp.status === true
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Distinct Until Changed (Response):&lt;/strong&gt;
Only update the DOM if the API response has actually changed. This reduces resource-heavy DOM updates. Use &lt;code&gt;distinctUntilChanged&lt;/code&gt; with a custom comparator function checking the &lt;code&gt;result_hash&lt;/code&gt;.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// ... (previous operators)
    .filter(resp =&gt; resp.status !== false)
    .distinctUntilChanged((a, b) =&gt; a.result_hash === b.result_hash)
// ...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This compares the &lt;code&gt;result_hash&lt;/code&gt; of the current and previous emissions. If they are the same, the data is filtered out.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You&apos;ve now created a pipeline that processes user input, performs checks, makes an API call, and formats the response, optimizing for resource usage.&lt;/p&gt;
&lt;h2 id=&quot;step-5--activating-everything-with-a-subscription&quot;&gt;Step 5 — Activating Everything with a Subscription&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;subscribe&lt;/code&gt; operator is the final link that connects an Observer to the Observable, enabling the flow of data. It typically implements three methods:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;onNext&lt;/code&gt;: Specifies what to do when an event is received.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;onError&lt;/code&gt;: Handles errors. No further &lt;code&gt;onNext&lt;/code&gt; or &lt;code&gt;onCompleted&lt;/code&gt; calls are made after an error.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;onCompleted&lt;/code&gt;: Called when &lt;code&gt;onNext&lt;/code&gt; has been called for the final time (no more data).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This enables lazy execution: the Observable pipeline is defined but only starts emitting data upon subscription.&lt;/p&gt;
&lt;p&gt;Subscribe to the Observable and route the data to the &lt;code&gt;showResults&lt;/code&gt; function:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-javascript&quot;&gt;// Full RxJS chain:
let searchInput = document.getElementById(&quot;search-input&quot;);

Rx.Observable.fromEvent(searchInput, &apos;input&apos;)
    .pluck(&apos;target&apos;, &apos;value&apos;)
    .filter(searchTerm =&gt; searchTerm.length &gt; 2)
    .debounceTime(500)
    .distinctUntilChanged()
    .switchMap(searchKey =&gt; Rx.Observable.ajax(`https://api.crossref.org/works?rows=50&amp;#x26;query.author=${searchKey}`)
        .map(resp =&gt; ({
            &quot;status&quot; : resp[&quot;status&quot;] == 200,
            &quot;details&quot; : resp[&quot;status&quot;] == 200 ? resp[&quot;response&quot;] : [],
            &quot;result_hash&quot;: Date.now()
        }))
    )
    .filter(resp =&gt; resp.status !== false)
    .distinctUntilChanged((a, b) =&gt; a.result_hash === b.result_hash)
    .subscribe(resp =&gt; showResults(resp.details)); // Pass only the details to showResults
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Save the &lt;code&gt;search-bar.html&lt;/code&gt; file. Open it in your web browser to test the search bar.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;(Image from the original tutorial: The completed search bar)&lt;/em&gt;
&lt;em&gt;(GIF from the original tutorial: Content being entered into the search bar, showing behavior with few characters)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;You have now subscribed to the Observable, activating your search bar.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;In this tutorial, you created a feature-rich search bar using RxJS, CSS, and HTML that provides real-time results. The search bar requires a minimum of three characters, updates automatically, and is optimized for both client and API server load.&lt;/p&gt;
&lt;p&gt;Complex requirements were addressed with concise RxJS code, leading to a solution that is reader-friendly and maintainable.&lt;/p&gt;
&lt;p&gt;For further reading, refer to the &lt;a href=&quot;http://reactivex.io/rxjs/identifiers.html&quot;&gt;official RxJS API documentation&lt;/a&gt;.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Generating Synthetic Music with JavaScript - Introduction to Neural Networks]]></title><description><![CDATA[Neural Networks (Image source: KOMA ZHANG — QUANTA MAGAZINE) Reference by Mayank Raj Follow Introduction Recurrent Neural Networks (RNN) are…]]></description><link>https://mayankraj.com/blog/synthetic-music-with-neural-networks</link><guid isPermaLink="false">https://mayankraj.com/blog/synthetic-music-with-neural-networks</guid><pubDate>Wed, 17 Apr 2019 19:40:46 GMT</pubDate><content:encoded>&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:4800/format:webp/1*lNTho6sG2ef0yovAB_n75g.jpeg&quot; alt=&quot;Neural Networks (Image source: KOMA ZHANG — QUANTA MAGAZINE)&quot;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://medium.com/cactus-techblog/generating-synthetic-music-with-javascript-introduction-to-neural-networks-a0b258fade40&quot;&gt;Reference&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;by &lt;a href=&quot;https://medium.com/@mayank9856?source=post_page---byline--a0b258fade40---------------------------------------&quot;&gt;Mayank Raj&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Follow&lt;/p&gt;
&lt;h2 id=&quot;introduction&quot;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Recurrent Neural Networks (RNN) are a way to consider the dimension of time when training or inferring from the Neural Network. While talking about Neural Networks in general it is assumed that inputs map to the same output irrespective of the order or sequence of inputs. This is true in cases when you have to identify objects in an image or predict a candidate’s eligibility for loan based on various circumstances. This is because each entity is independent of the previous entity. But Neural Networks built on this assumption do not perform well in cases where the input sequence is one of the most crucial features like machine translation where a sequence of words matter, or in case of system log analysis where the order of events matter among other things.&lt;/p&gt;
&lt;p&gt;In this tutorial we will briefly understand how LSTM considers a certain sequence of events. We will then explore &lt;a href=&quot;https://magenta.tensorflow.org/&quot;&gt;Magenta&lt;/a&gt; &lt;a href=&quot;https://magenta.tensorflow.org/&quot;&gt;a&lt;/a&gt;, an ML library that can be used to generate music and art. We will also build a small app that will play a different sequence of drums for us each time by using &lt;a href=&quot;https://github.com/tensorflow/magenta/tree/master/magenta/models/drums_rnn&quot;&gt;DrumRNN&lt;/a&gt; in the browser using &lt;a href=&quot;https://tensorflow.github.io/magenta-js/music/&quot;&gt;Magenta.js&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;step-1-understanding-neural-networks-and-recurrent-neural-networks&quot;&gt;Step 1: Understanding Neural Networks and Recurrent Neural Networks&lt;/h2&gt;
&lt;p&gt;Neural Networks are universal function approximators, i.e.given an adequate size of neural networks you can use it to define any function. In it’s simplest form neural network can be understood as a Perceptron. It has the following components:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;&lt;em&gt;Inputs&lt;/em&gt;:&lt;/strong&gt; The inputs to the perceptron&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;em&gt;Weights&lt;/em&gt;:&lt;/strong&gt; These are the weights assigned to each input&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;em&gt;Bias&lt;/em&gt;:&lt;/strong&gt; This is a constant that is added to the inputs&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;em&gt;Activation function&lt;/em&gt;:&lt;/strong&gt; It maps input to output&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:1400/format:webp/1*xf4PpCB6p138H_-wi2Qqdw.png&quot; alt=&quot;Perceptron (Image source: https://www.researchgate.net/publication/327392288_A_Quantum_Model_for_Multilayer_Perceptron)&quot;&gt;&lt;/p&gt;
&lt;p&gt;Weights help to give importance to certain inputs more than others. For instance, in order to decide if John should be granted a loan or not, the input related to his income should have a higher weight compared to the one that represents the initials of his name. Bias helps to make sure that given the same input, different perceptrons get different summed input by adding or subtracting a constant value from the input. This summed input is then passed through an activation function. This function has one job - to map the input to the output. There are many variations of activation functions, each suited for a different job. You would use a sigmoid activation function when you want to scale values in the range of 0 to 1, which is generally the case in a multi-class classification. In cases of binary classification where you want to know just one of the two classes- 0 or 1, you would use something like a binary step.&lt;/p&gt;
&lt;p&gt;In essence, this perceptron is also a single-layer neural network. When you stack number of such perceptrons in a layer and increase the number of layers you get a deep neural network. The training data is used to adjust the parameters of each such perceptron in this network such as weights, biases etc. As a convention, the layer that receives the input is also known as the input layer and the one that provides output is called the output layer. All the layers in between, if present are known as hidden layers.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:1400/format:webp/1*kaxFfBnq9MvXQQ_b8alIxA.png&quot; alt=&quot;Neural Network (Image source: https://en.wikipedia.org/wiki/Artificial_neural_network)&quot;&gt;&lt;/p&gt;
&lt;p&gt;As you might have observed before, the output of a perceptron is only dependent on its input at a certain time. As they are the building blocks of a neural network, the output of the neural network is also dependent on only the input at a certain time. This is where recurrent neural networks come into the picture.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://miro.medium.com/v2/resize:fit:458/format:webp/1*vUfazfSJwrhklcpsUd9qDA.png&quot; alt=&quot;Recurrent neural networks&quot;&gt;&lt;/p&gt;
&lt;p&gt;The output from the current state is fed again during the next state viz. the hidden layer. This makes the network capable of knowing what happened in the previous state or when the last input was passed through the network. As the previous state also has information about the states before it, this input represents the complete history to some extent. The term - “some extend” is used as there is a concept of a vanishing gradient which makes the memory of a state shrink as and when new states are seen.&lt;/p&gt;
&lt;p&gt;In this section you leaned about what makes up a neural network and how they work. You also saw how recurrent neural networks are different from normal neural network and also how they work with time series data.&lt;/p&gt;
&lt;h2 id=&quot;step-2-introduction-to-tensorflowjs-to-magentajs&quot;&gt;Step 2: Introduction to TensorFlow.js to Magenta.js&lt;/h2&gt;
&lt;p&gt;In this section you will learn about the Magenta project and look at Magenta.js. Although training the models is out of scope of this tutorial you will see how to use pretrained models to generate synthetic music using only JavaScript.&lt;/p&gt;
&lt;p&gt;Magenta is a project lead by Google that focuses on using neural network in the domain of music and art. Although the core concepts can be implemented with any machine leaning library, the team behind Magenta has used TensorFlow for the task. TensorFlow is a library for machine learning developed and used by Google. It is an end-to-end system which means it is easy to not only develop with TensorFlow by using GPU’s and TPU’s for training but also deploy with it to multiple machines and even IOT devices. It also has TensorFlow.js which makes it easy to develop models and infer from them all within the JavaScript ecosystem. This means you can train and use the model in the browser itself.&lt;/p&gt;
&lt;p&gt;In the next section you will use Magenta.js to create synthetic music.&lt;/p&gt;
&lt;h2 id=&quot;step-3-creating-the-base&quot;&gt;Step 3: Creating the base&lt;/h2&gt;
&lt;p&gt;In this section we will design the base of the app. We will load all the libraries we need using Content Delivery Network or CDN. Start by creating a file named &lt;code&gt;index.html&lt;/code&gt; and write the following snippet to it:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;[index.html]
&amp;#x3C;html&gt;
&amp;#x3C;head&gt;
    &amp;#x3C;title&gt;Synthetic Music&amp;#x3C;/title&gt;
    &amp;#x3C;style type=&quot;text/css&quot;&gt;
        body {
            margin: 0;
            padding: 0;
            background-color: #f6f6f6;
            font-family: sans-serif;
        }        .play {
            width: fit-content;
            margin: 0 auto;
            padding: 12px;
            border: 2px solid #323232;
            border-radius: 50%;
        }        .play div {
            border-top: 10px solid transparent;
            border-bottom: 10px solid transparent;
            border-left: 20px solid #323232;
            height: 0px;
            width: fit-content;
            margin: 0 auto;
        }        .heading {
            text-align: center;
            margin-top: 5rem;
        }        #pattern-container {
            margin-top: 2rem;
            width: fit-content;
            border-radius: 5px;
            margin: 0 auto;
        }        .pattern-group {
            display: inline-block;
            background-color: #e4f9f5;
            padding: 5px 6px;
        }        .pattern-group.seed {
            background-color: #a8e6cf;
        }        .pattern.active {
            background-color: #11999e;
        }        .pattern {
            height: 10px;
            width: 5px;
            display: block;
            margin: 5px 0;
            padding: 3px;
            border-radius: 4px;
            color: transparent;
            background-color: #cbf1f5;
        }
    &amp;#x3C;/style&gt;
&amp;#x3C;/head&gt;&amp;#x3C;body&gt;
&amp;#x3C;h2 class=&quot;heading&quot;&gt;Synthetic Music with Neural Networks &amp;#x3C;/h2&gt;
&amp;#x3C;div class=&quot;play&quot; onclick=&quot;createAndPlayPattern(this)&quot;&gt;
    &amp;#x3C;div&gt;&amp;#x3C;/div&gt;
&amp;#x3C;/div&gt;
&amp;#x3C;div&gt;
    &amp;#x3C;div id=&apos;pattern-container&apos;&gt;&amp;#x3C;/div&gt;
&amp;#x3C;/div&gt;
&amp;#x3C;script src=&apos;https://code.jquery.com/jquery-3.3.1.slim.min.js&apos;&gt;&amp;#x3C;/script&gt;
&amp;#x3C;script type=&apos;text/javascript&apos; src=&apos;https://cdn.jsdelivr.net/npm/lodash@4.17.5/lodash.min.js&apos;&gt;&amp;#x3C;/script&gt;
&amp;#x3C;script type=&apos;text/javascript&apos; src=&apos;https://gogul09.github.io/js/tone.js&apos;&gt;&amp;#x3C;/script&gt;
&amp;#x3C;script type=&apos;text/javascript&apos;
        src=&apos;https://cdn.jsdelivr.net/npm/@magenta/music@0.0.8/dist/magentamusic.min.js&apos;&gt;&amp;#x3C;/script&gt;
&amp;#x3C;script&gt;
    // Continue here...
&amp;#x3C;/script&gt;
&amp;#x3C;/body&gt;
&amp;#x3C;/html&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We have defined a few custom CSS styles in the head section. Additional DOM elements were also defined to which we will later add the music patterns generated by out neural network. We have also loaded a few libraries like &lt;code&gt;jQuery&lt;/code&gt; and &lt;code&gt;loadash&lt;/code&gt; to clean up our code by using the API provided by them, &lt;code&gt;tone.js&lt;/code&gt; to play tones and finally &lt;code&gt;magenta.js&lt;/code&gt; which we will use to generate music patterns.&lt;/p&gt;
&lt;h2 id=&quot;step-4-setting-up-tonejs&quot;&gt;Step 4: Setting up Tone.js&lt;/h2&gt;
&lt;p&gt;In this section we will setup &lt;code&gt;Tone.js&lt;/code&gt; so that it knows what notes are available to play. We will be playing drums, so we need sounds for different pieces of drums. You can get these sound files from the assets folder in this &lt;a href=&quot;https://github.com/rajmayank/synthetic-music-with-neural-networks&quot;&gt;repository&lt;/a&gt;. You will find files named &lt;code&gt;808-hihat-open-vh.mp3&lt;/code&gt;, &lt;code&gt;808-hihat-vh.mp3&lt;/code&gt;, &lt;code&gt;808-kick-vh.mp3&lt;/code&gt;, &lt;code&gt;909-clap-vh.mp3&lt;/code&gt;, &lt;code&gt;909-rim-vh.wav&lt;/code&gt;, &lt;code&gt;flares-snare-vh.mp3&lt;/code&gt;, &lt;code&gt;slamdam-tom-high-vh.mp3&lt;/code&gt;, &lt;code&gt;slamdam-tom-low-vh.mp3&lt;/code&gt;, &lt;code&gt;slamdam-tom-mid-vh.mp3&lt;/code&gt; and &lt;code&gt;small-drum-room.wav&lt;/code&gt;. Among these &lt;code&gt;small-drum-room.wav&lt;/code&gt; will be used to create reverb.&lt;/p&gt;
&lt;p&gt;Start by creating a convolver which you will later use to create reverb. Also make sure that the wet value is set to 0.3 which means 30% of this effect will be applied to the tone on which it is used. Write the following snippet to do so.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;let reverb = new Tone.Convolver(`assets/small-drum-room.wav`).toMaster();
reverb.wet.value = 0.3;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next we have to define the individual components of the drum kit. These are Kick, Snare, Hi-hat closed, Hi-hat open, Tom low, Tom mid, Tom high, Clap and Rim. Write the following snippet that creates a drumkit with these components in order.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;let drumKit = [    new Tone.Player(`assets/808-kick-vh.mp3`).toMaster(),
    new Tone.Player(`assets/flares-snare-vh.mp3`).toMaster(),
    new Tone.Player(`assets/808-hihat-vh.mp3`).connect(new Tone.Panner(-0.5).connect(reverb)),
    new Tone.Player(`assets/808-hihat-open-vh.mp3`).connect(new Tone.Panner(-0.5).connect(reverb)),
    new Tone.Player(`assets/slamdam-tom-low-vh.mp3`).connect(new Tone.Panner(-0.4).connect(reverb)),
    new Tone.Player(`assets/slamdam-tom-mid-vh.mp3`).connect(reverb),
    new Tone.Player(`assets/slamdam-tom-high-vh.mp3`).connect(new Tone.Panner(0.4).connect(reverb)),
    new Tone.Player(`assets/909-clap-vh.mp3`).connect(new Tone.Panner(0.5).connect(reverb)),
    new Tone.Player(`assets/909-rim-vh.wav`).connect(new Tone.Panner(0.5).connect(reverb))
];
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Notice how each component is defined as an instance of Tone Player. Components that need reverb effect applied to them, a tone panner was connected to the original tone player. Panning is a process of placing the instrument in the 3D space by deciding the shape of the signal that is given to each channel of audio. Suppose, in a 2 channel the audio, you pass a tone signal equally through both the channels, that creates the illusion of the tone originating from the centre, you can then move this perceived source of the tone by controlling the signal that goes to each of the channel.&lt;/p&gt;
&lt;p&gt;In this section you created the components that make up your drumkit. You also loaded the sound files for each tone and added reverbs to the tones that require it.&lt;/p&gt;
&lt;h2 id=&quot;setting-up-magentajs&quot;&gt;Setting up Magenta.js&lt;/h2&gt;
&lt;p&gt;In this section you will use neural network to continue a provided seed pattern and create synthetic music. Before we actually start using the neural networks these are a few methods that need to be setup that will convert the music notes to sequences that the model understands. Similarly, we will have to convert back the sequence that the model gave back to a tone that can be played.&lt;/p&gt;
&lt;p&gt;Start by creating a mapping of MIDI values for each tone. MIDI or Musical Instrument Digital Interface is the protocol by which various devices that deal with music communicate with each other. If you connect some synthesizer with your computer, this is the interface used for communication. As you may have guessed this is also how we will communicate with our model. Magenta has been designed in this way as it makes it a easy plug in existing devices into the model directly. You may have guessed it, there are already &lt;a href=&quot;https://nsynthsuper.withgoogle.com&quot;&gt;devices available&lt;/a&gt; that do just that.&lt;/p&gt;
&lt;p&gt;A machine leaning model performs best when the range of input is defined. Thus we will create a mapping of values from MIDI range into the range that the model can work with. Use the following variables to create this mapping:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;const midiDrums = [36, 38, 42, 46, 41, 43, 45, 49, 51];
const reverseMidiMapping = new Map([    [36, 0], [35, 0], [38, 1], [27, 1], [28, 1], [31, 1], [32, 1], [33, 1],
    [34, 1], [37, 1], [39, 1], [40, 1], [56, 1], [65, 1], [66, 1], [75, 1],
    [85, 1], [42, 2], [44, 2], [54, 2], [68, 2], [69, 2], [70, 2], [71, 2],
    [73, 2], [78, 2], [80, 2], [46, 3], [67, 3], [72, 3], [74, 3], [79, 3],
    [81, 3], [45, 4], [29, 4], [41, 4], [61, 4], [64, 4], [84, 4], [48, 5],
    [47, 5], [60, 5], [63, 5], [77, 5], [86, 5], [87, 5], [50, 6], [30, 6],
    [43, 6], [62, 6], [76, 6], [83, 6], [49, 7], [55, 7], [57, 7], [58, 7],
    [51, 8], [52, 8], [53, 8], [59, 8], [82, 8]]);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Next, also set the temperature of the music you desire and also the length of the pattern as follows:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;const temperature = 1.0;
const patternLength = 32;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To convert the notes into a sequence, write down the following function:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function fromNoteSequence(seq, patternLength) {
    let res = _.times(patternLength, () =&gt; []);
    for (let { pitch, quantizedStartStep } of seq.notes) {
        res[quantizedStartStep].push(reverseMidiMapping.get(pitch));
    }
    return res;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It takes two inputs, the sequence of notes and the pattern length. Loadash was used to first creat a list of length defined by &lt;code&gt;patternLength&lt;/code&gt; consisting of empty lists that then filled up next. Individual notes are then iterated over to unpack the start time of the note and then the note is placed at that interval by using the &lt;code&gt;reverseMidiMapping&lt;/code&gt; defined earlier.&lt;/p&gt;
&lt;p&gt;The output from model can now be converted into a sequence that can be used to play the drum kit. The reverse has to happen as well i.e. pattern sequences should be converted into note sequences that the model can understand. Write down the following function that performs this job:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function toNoteSequence(pattern) {
        return mm.sequences.quantizeNoteSequence({
                ticksPerQuarter: 220,
                totalTime: pattern.length / 2,
                timeSignatures: [{
                    time: 0,
                    numerator: 4,
                    denominator: 4
                }],
                tempos: [{
                    time: 0,
                    qpm: 120
                }],
                notes: _.flatMap(pattern, (step, index) =&gt;
                    step.map(d =&gt; ({
                        pitch: midiDrums[d],
                        startTime: index * 0.5,
                        endTime: (index + 1) * 0.5
                    }))
                )
            },
            1
        );
    };
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here are few details on the attributes that have been defined when creating this note sequence:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;ticksPerQuarter&lt;/code&gt;: &lt;code&gt;ticks&lt;/code&gt; is the unit of time in the MIDI standard.
&lt;code&gt;totalTime&lt;/code&gt;: The length of the sequence that has been provided in the &lt;code&gt;notes&lt;/code&gt; attribute. This length is measured in terms of quantized steps.
&lt;code&gt;timeSignatures&lt;/code&gt;: This defines the time signature used in musical notation.
&lt;code&gt;tempos&lt;/code&gt;: Define the temps used in the tone sequence provided. &lt;code&gt;qpm&lt;/code&gt; here refers to quarter notes per minute.
&lt;code&gt;notes&lt;/code&gt;: This represents the notes with the pitch and duration of each note in the sequence.&lt;/p&gt;
&lt;p&gt;Once we have the pattern, there should be a way to play it as well. To play the pattern we will use the drumkit we created in the previous section. Write the following method that takes a pattern and plays it using Tone.js:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function playPattern(pattern) {
    sequence = new Tone.Sequence(
        (time, {drums, index}) =&gt; {
            drums.forEach(d =&gt; {
                drumKit[d].start(time)
            });
        },
        pattern.map((drums, index) =&gt; ({ drums, index })),
        &apos;16n&apos;
    );    Tone.context.resume();
    Tone.Transport.start();
    sequence.start();
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;Tone.Sequence&lt;/code&gt; has been used to accomplish this. It expects three inputs: &lt;code&gt;callback&lt;/code&gt;: This method would be called for each event
&lt;code&gt;event&lt;/code&gt;: The individual events of this sequence. Here we map each individual pattern as an object that has the drum components to play as &lt;code&gt;drums&lt;/code&gt; and the index of the pattern in the sequence as &lt;code&gt;stepId&lt;/code&gt;
&lt;code&gt;subdivision&lt;/code&gt;: The subdivision between the events &lt;code&gt;Tone.context.resume&lt;/code&gt; then starts the audio context which is required to connect to the audio interface provided by the browser.
&lt;code&gt;Tone.Transport&lt;/code&gt; makes sure the timing of the music stays perfect by not directly relying on browsers timing. &lt;code&gt;sequence.start&lt;/code&gt; finally sets everything in motion.&lt;/p&gt;
&lt;p&gt;With this you are now ready to finally use the model to create music. Before we actually start using the model let us look at how note pattern look. A specific pattern in a pattern sequence is an array of indexes from 0 to 8 of the 9 components in our drumkit. So a pattern &lt;code&gt;[0,2,4]&lt;/code&gt; would play the Kick, Hi-hat closed and Tom low.&lt;/p&gt;
&lt;p&gt;We now use a seed pattern that our model will then improvise upon. Use the following seed pattern for it.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;var seedPattern = [    [0, 2],
    [0],
    [2, 5, 8],
    [],
    [2, 5, 8],
    [],
    [0, 2, 5, 8],
    [4, 5, 8],
    [],
    [0, 5, 8]
];
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Use the following snippet to define a function that creates this seed pattern and then creates a pattern of length &lt;code&gt;patternLength&lt;/code&gt; which was previously set to &lt;code&gt;32&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;let drumRnn = new mm.MusicRNN(&apos;https://storage.googleapis.com/download.magenta.tensorflow.org/tfjs_checkpoints/music_rnn/drum_kit_rnn&apos;);
drumRnn.initialize();function createAndPlayPattern() {
    drumRnn
        .continueSequence(seedSeq, patternLength, temperature)
        .then(r =&gt; seedPattern.concat(fromNoteSequence(r, patternLength)))
        .then(displayPattern)
        .then(playPattern)
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We start by first loading the Drum RNN model and then initializing it. We have then created a function that uses this model to take a seed pattern and then continue it to create a new pattern using the &lt;code&gt;continueSequence&lt;/code&gt; method. We then pipe this note sequence to convert it to sequence that we can play our drumkits with. We have also chained the &lt;code&gt;displayPattern&lt;/code&gt; to visualize the pattern and then finally play the pattern with &lt;code&gt;playPattern&lt;/code&gt; method. We will create the visualization method in the next section.&lt;/p&gt;
&lt;h2 id=&quot;adding-visualization&quot;&gt;Adding visualization&lt;/h2&gt;
&lt;p&gt;In this section we will toe up everything to create visualization for our pattern. Modern browsers like Chrome require the user to make some interaction on the page before it allows the music to be played. This is actually a usability feature. You would not like it if a pop-up that opened in the background suddenly starts playing something. To make sure that we start our process only after the user interacts with the page, we had created a play button when writing the HTML and added an &lt;code&gt;onclick&lt;/code&gt; attribute to it. The method referenced in this attribute was &lt;code&gt;createAndPlayPattern&lt;/code&gt; which we created in the previous step that takes a seed pattern, generates music, visualizes it and plays it. We do not want this play button to be visible after the first click. Modify the &lt;code&gt;createAndPlayPattern&lt;/code&gt; pattern to accept the click event and delete the element from the page as follows:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function createAndPlayPattern(e) {
    $(e).remove()
    let seedSeq = toNoteSequence(seedPattern);
    drumRnn
        .continueSequence(seedSeq, patternLength, temperature)
        .then(r =&gt; seedPattern.concat(fromNoteSequence(r, patternLength)))
        .then(displayPattern)
        .then(playPattern)
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;You are now ready to add the last and final part to the puzzle - &lt;code&gt;displayPattern&lt;/code&gt; method that visualizes the pattern. Write down the following method to achieve this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;function displayPattern(patterns) {
    for (let patternIndex = 0; patternIndex &amp;#x3C; patterns.length; patternIndex++) {
        let pattern = patterns[patternIndex];
        patternBtnGroup = $(&apos;&amp;#x3C;div&gt;&amp;#x3C;/div&gt;&apos;).addClass(&apos;pattern-group&apos;);
        if (patternIndex &amp;#x3C;= seedPattern.length)
            patternBtnGroup.addClass(&apos;seed&apos;);        for (let i = 0; i &amp;#x3C;= 8; i++) {
            if (pattern.includes(i))
                patternBtnGroup.append($(`&amp;#x3C;span&gt;&amp;#x3C;/span&gt;`).addClass(&apos;pattern active&apos;));
            else
                patternBtnGroup.append($(`&amp;#x3C;span&gt;&amp;#x3C;/span&gt;`).addClass(&apos;pattern&apos;));
        }
        $(&apos;#pattern-container&apos;).append(patternBtnGroup)
    }
    return patterns;
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It takes the pattern and creates multiple spans on the page. Each span is sorted out in a vertical stack of 9 blocks representing each component of the drum kit. It is highlighted if it is on or played when the respective sequence is played. Also the seed pattern is highlighted to make it easy to see what the model has generated.&lt;/p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Recurrent neural networks provide a mechanism to retain state over input. This makes RNN a great fit to predict music as at it’s core, music is a sequence of tones which are all dependent on each other. Magenta is a great starting point for any such projects.&lt;/p&gt;
&lt;p&gt;Check out the source code at &lt;a href=&quot;https://github.com/rajmayank/synthetic-music-with-neural-networks&quot;&gt;github&lt;/a&gt; and a live demo &lt;a href=&quot;https://rajmayank.github.io/synthetic-music-with-neural-networks/&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;</content:encoded></item><item><title><![CDATA[Automating windows with COM object]]></title><description><![CDATA[A few weeks ago I got a project wherein I had to get various statistical data about a document. One of the important parameters was to get a…]]></description><link>https://mayankraj.com/blog/automation-with-windows-com</link><guid isPermaLink="false">https://mayankraj.com/blog/automation-with-windows-com</guid><pubDate>Sat, 01 Dec 2018 16:54:41 GMT</pubDate><content:encoded>&lt;p&gt;A few weeks ago I got a project wherein I had to get various statistical data about a document. One of the important parameters was to get a precise count of grammatical errors in the document. I tried to go the open source way and tested out &lt;a href=&quot;https://www.languagetool.org/&quot;&gt;LanguageTool&lt;/a&gt;, &lt;a href=&quot;http://proselint.com/&quot;&gt;PorseLink&lt;/a&gt; among others. None of these were up to the mark of the good old MS-Word.
So I had to now work on a way to problematically access Word, load a document in it and get the statistics.&lt;/p&gt;
&lt;p&gt;Python being versatile and great at memory management, we would be using it to control COM object and do what we need to do.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;This is the Part 1 of the two series post.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Part 1&lt;/strong&gt; : Introduction to &lt;code&gt;COMobject&lt;/code&gt;, initialization, examples etc&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Part 2&lt;/strong&gt; : Setting up a server environment. Making use of threads and opening using thread for each process. This is how to communicate with the &lt;code&gt;COMobject&lt;/code&gt; in a multi-thread environment as it cannot be passed to a thread directly.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;before-we-dig-deep-what-is-comobject-&quot;&gt;Before we dig deep, what is COMobject ?&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;The Microsoft Component Object Model (COM) is a platform-independent, distributed, object-oriented system for creating binary software components that can interact. COM is the foundation technology for Microsoft&apos;s OLE (compound documents), ActiveX (Internet-enabled components), as well as others. &lt;a href=&quot;https://msdn.microsoft.com/en-us/library/windows/desktop/ms694363(v=vs.85).aspx&quot;&gt;read more...&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In a nutshell it can be used to control various aspects of Windows. We would be focusing on using the &lt;code&gt;COMobject&lt;/code&gt; to control MS-Office applications like Word and Excel. Although it is not limited to just MS-Office applications, it can be used to communicate with other softwares like IE (if you&apos;re still into it...)&lt;/p&gt;
&lt;p&gt;You can do things like :&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Create triggers to perform various tasks
&lt;ul&gt;
&lt;li&gt;Add new rows to an excel sheet in real time with the current scores of a game. This excel sheet can then have some complex algorithms to predict which team would win by performing calculations on all the rows. Although it can be done with implementation in even NodeJS but we don&apos;t want to take the pain to port the functions/macros from Excel to JS&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Log the current changes made by the user XYZ in the database&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The possibilities are limitless. I guess you already have a use case in mind and landing to this page to know the process, if not it&apos;s always good to know if there is possibility to get a certain functionality done.&lt;/p&gt;
&lt;h2 id=&quot;python-and-com&quot;&gt;Python and COM&lt;/h2&gt;
&lt;p&gt;We will need the following tools to communicate with MS-Word :&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;MS Office or Specific applications already installed&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://pypi.python.org/pypi/pywin32&quot;&gt;PyWin32 library&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;code&gt;PyWin32&lt;/code&gt; is a great library that gives us the same set of methods/properties that are exposed by COM which are natively in Visual Basic for Applications (VBA).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;NOTE: I&apos;m using Python3 in the code samples in this guide.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;lets-start&quot;&gt;Let&apos;s start&lt;/h2&gt;
&lt;p&gt;To get hold of a COM object&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;import win32com.client

# for MS-Word
word = win32.gencache.EnsureDispatch(&apos;Word.Application&apos;)
# OR
word = win32com.client.DispatchEx(&apos;Word.Application&apos;)

# for MS-Excel
excel = win32.gencache.EnsureDispatch(&apos;Excel.Application&apos;)
# OR
excel = win32com.client.DispatchEx(&apos;Excel.Application&apos;)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;There are some key differences to be noted here&lt;/p&gt;
&lt;p&gt;&lt;code&gt;EnsureDispatch&lt;/code&gt; : In dead simple terms, it makes sure that a object which references the specified application is returned. If the application is already open, it will return the same instance, if not it will start a instance and return it.
It is useful in cases where you would want to make use of single application that holds certain meta-data, over the life of the process.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;DispatchEx&lt;/code&gt; : For each call, you would get a new object returned. This would mean that if you call it 7 times, you would see 7 instances of MS-Word in task manager.
It makes sense to get independent instances when you would want to handle each process independently and close the application when one process completes.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Understand that if you run a web-server that takes in data, puts it in an Excel, runs a macro and return a result. You would want to have a dedicated instance of Excel for each process so that you can close it as soon as the response is ready. If you used &lt;code&gt;EnsureDispatch&lt;/code&gt; or &lt;code&gt;Dispatch&lt;/code&gt;, that would mean that at the end of each process, you would have to check if any other process is in progress, if not shut down the application.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;It&apos;s requires a bit extra memory but helps to keep memory leak in check as using same application instance to load up multiple files that do not require each others data, means that with each load, some meta-data is loaded which is not cleared when closing the process and it builds up over time. &lt;em&gt;We will cover this aspect in detail in the Part-2 of the post.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Coming back to the point, let&apos;s go ahead.&lt;/p&gt;
&lt;p&gt;Now that we have the &lt;code&gt;COMobject&lt;/code&gt; we can use it a open a Word Document.&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;doc = word.Documents.Open( os.path.join( os.getcwd(), &apos;path/to/files/&apos;, filename), ReadOnly=True)
&lt;/code&gt;&lt;/pre&gt;
&lt;blockquote&gt;
&lt;p&gt;NOTE : As you might have had guessed, you can use the same &lt;code&gt;word&lt;/code&gt; object to open multiple &lt;code&gt;doc&lt;/code&gt; objects. This is what I was referring to using same application to open multiple documents in the last paragraph&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;
&lt;p&gt;At this point, we can refer to Microsoft&apos;s documentation to get gist of various properties &amp;#x26; functions available.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://msdn.microsoft.com/en-us/library/bb244391(v=office.12).aspx&quot;&gt;&lt;strong&gt;Word 2007 Developers Reference&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://msdn.microsoft.com/en-us/library/bb244515(v=office.12).aspx&quot;&gt;&lt;strong&gt;Word Object Model reference&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://msdn.microsoft.com/en-us/library/bb244569(v=office.12).aspx&quot;&gt;&lt;strong&gt;Application Object&lt;/strong&gt;&lt;/a&gt; : List of methods and properties available&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;lets-have-look-at-some-usage-examples&quot;&gt;Let&apos;s have look at some usage examples:&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Make the application window invisible :&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;word.Visible = False
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;To add a new Document&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;doc = word.Documents.Add()
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;To get all the grammatical or spelling errors&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;grammaticalErrors = doc.GrammaticalErrors
spellingErrors = doc.doc.SpellingErrors

#or to simple get the count
grammaticalErrorsCount = doc.GrammaticalErrors.Count
spellingErrorsCount = doc.SpellingErrors.Count
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;If you have a look at the &lt;a href=&quot;https://msdn.microsoft.com/en-us/library/bb237490(v=office.12).aspx&quot;&gt;&lt;strong&gt;Word BuiltInProperty&lt;/strong&gt;&lt;/a&gt;, you&apos;ll see you can do the following&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;wordCount = int(doc.BuiltInDocumentProperties(15))
characterCount = int(doc.BuiltInDocumentProperties(16))
paragraphCount = int(doc.BuiltInDocumentProperties(24))
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;or do something much more complex as&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;characterPerWord = doc.ReadabilityStatistics(&quot;Characters per Word&quot;).Value
&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;or something as simple as&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;shapeCount = doc.Shapes.Count
tableCount = doc.Tables.Count
sentenceCount = doc.Sentences.Count
tableOfContentCount = doc.TablesOfContents.Count

#loop through each paragraphs
for paras in doc.Paragraphs:
    # paras will now be each individual paragraph

&lt;/code&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;As you can see the possibilities are endless. MS-Office tools are pretty powerful in what they do. We did not even dig into the excel territory. Majority of the questions on stackoverflow are based on excel as automating tasks there makes can be of great use.&lt;/p&gt;
&lt;p&gt;In the next post we will focus on how to wrap all this up into dedicated thread for each process and word with COMobject in threads.&lt;/p&gt;</content:encoded></item></channel></rss>