About Me

My photo
Rohit is an investor, startup advisor and an Application Modernization Scale Specialist working at Google.

Saturday, January 25, 2020

Application Portfolio Rationalization for Cloud Migration

How to run a an workload/app migration discovery workshop 

  1. Goals (Technical and Business) for the program > (Objectives & Key Results) for the session
  2. Start at the portfolio level. Figure out how may high level portfolio’s exist.
  3. For each portfolio figure out the high level buckets of apps. case in point - JavaEE apps, .NET apps, Spring Apps, PHP/Python/NodeJS …. you should create a heat map of the portfolio with the buckets. Each tile in the heat map represents a bucket and the color + intensity shows the ease of migration. The size of the tile represents the % in the portfolio. 
  4. Now you have two choices go broad or go deep. i.e. look at one bucket and dig deeper or go broad and sample a couple from each bucket.
  5. Settle on a couple of specific apps to start with to drill down throughout the process
    Apply business value and other org heuristics to prioritize the apps in terms of business ROI. 2 * 2 matrix of technical effort vs business value to determine focus. This can be done at the bucket level or individual app level. For individual apps run a brief SNAP to get an idea of cloud suitability. Customized SNAP with Heat maps is the way to go … so low tech way of doing this is to do this on a flipchart with Blue grid easel pads like or go with excel
  6. What is the smallest possible thing we could do to add value. Discussion of MVPs. What will a potential AppTx engagement look like. How will we measure success and elicit feedback.
  7. If they you are stuck in the paralysis part then do a path to production or value stream exercise to figure out what is the top constraint/problem you need to focus ON in addition to the apps
  8. Retrospective & Next Steps

credit to Felicia Schwartz who developed this model ...


Bucketing and Technical Suitability of Applications
Application Migration Heatmap

Thursday, January 16, 2020

Rohit Kelapure A Year In Review 2019

So this is a bit late - but there is never a better time to retrospect and reflect. Without the support of my awesome AppTx team, peers and management at Pivotal this would not have been possible.

Rohit Kelapure - A Year In Review 2019

# Delivery

I  anchored seven enterprise engagements including one that was featured in the Wall Street Journal. I led an finished initiatives in a fifty person solution architect team in the following areas > Training & Cross-functional Enablement, Practice Management, Internal Initiatives, Tooling, Scoping, Selling, Recruiting, Marketing, Blog Posts, Webinars and Conference sessions. See details below.

# Practice Management 

  • Cookbooks Maintenance
  • Created Healthcheck & SRE Offering 
  • App Services Anchor Bootcamp
  • Blog Series Architecture - A Pivotal Opinion (upcoming)
  • Mainframe Modernization GTM
  • Closing loops - Feedback From AppTx to R&D
  • Closing loops - Feedback to DATA, PCFMetrics, Spring Cloud Services and other R&D Teams
  • Active participation PWG-SRE practice workgroup
  • AppTx for PSR Offering

# Training & Cross-functional Enablement 

  • Kubernetes Training mini-Conference
  • So You Want To Run An AppTx Scoping
  • So You Want To Run An AppTx Healthcheck
  • Anchoring Best Practices - Things We have learnt the hard way
  • Wrote On-boarding 2 - Week - New Hire Fast Ramp 4 AppTx Solution Architects
  • Continuous assistance on Slack #modern-family & #app-transformation channels
  • AppTx Scoping Retro - Train the scopers 2 sessions 
  • PAL PKS Course Development 
  • Microservices Workshop DBSBank

# Pivotal Internal Initiatives

  • Google Anthos intel
  • PAA-RFC
  • Cookbooks - RFC
  • Java Devex Team

# Created Tools

  1. Pivotal App Analyzer  _Refinement of Rules , Guidance_
  2. AppTx Effort Estimation Model with Steve Woods
  3. PKS SNAP
  4. Spring Bootifier with Tim Dalsing

# Recruiting - 2 Solution Architects

# Mentored/Improved - One colleague

# Conducted Scopings - 6 including commercial and federal 

# Pivotal Blog Posts

- Twitter  @rkela (750 followers) 

# Webinars (Solo, Partners & Customers)

  1. Why Your Digital Transformation Strategy Demands Middleware Modernization
  2. How to Migrate Applications Off a Mainframe
  3. Tools and Recipes to Replatform Monolithic Apps to Modern Cloud Environments 
  4. App Modernization with .NET Core: How Travelers Insurance is Going Cloud-Native 

# Conferences - SpringOne Platform 2019

- [360-Degree Health Assessment of Microservices on the PCF Platform](https://springoneplatform.io/2019/sessions/360-degree-health-assessment-of-microservices-on-the-pcf-platform)

# Certifications


# Courses


Wednesday, January 15, 2020

Create SRE Dashboards For Applications


SLI/SLO Dashboards are a MUST for every SRE team as a core practice to manage end user expectations and needs. Here is a neat example of a SLO dashboard for the Pivotal Platform



So how do you craft these SLO dashboards. There are some APM tools that are ahead of the game when it comes to SLI/SLO dashboard templates.  You can customize the stock template to build your dashboard. See links below  ....

HoneyComb


DataDog

New Relic


https://docs.newrelic.com/docs/using-new-relic/welcome-new-relic/measure-devops-success/establish-objectives-baselines-define-team-slos

https://docs.newrelic.com/docs/browser/new-relic-browser/getting-started/introduction-new-relic-browser

https://docs.newrelic.com/docs/using-new-relic/welcome-new-relic/measure-devops-success/establish-objectives-baselines-define-team-slos

https://blog.newrelic.com/engineering/programmatic-service-level-indicator/

https://blog.newrelic.com/product-news/most-popular-new-relic-one-applications-4/

https://rpm.newrelic.com/accounts/2589077/setup

https://github.com/newrelic/nr1-slo-r#creating-a-webhook-to-forward-alert-incidents-to-insights



Tuesday, December 3, 2019

Interesting :aws: Reinvent announcements threads for app-modernization

Amazon EventBridge schema registry stores event structure - or schema - in a shared central location and maps those schemas to code for Java, Python, and Typescript so it’s easy to use events as objects in your code.  https://aws.amazon.com/about-aws/whats-new/2019/12/introducing-amazon-eventbridge-schema-registry-now-in-preview/?trk=ls_card

AWS launches new program to drive migrations for end of support Windows Server applications https://aws.amazon.com/about-aws/whats-new/2019/12/aws-launches-program-drive-migration-windows-server/?trk=ls_card

Amazon Managed Apache Cassandra Service - Eat Databricks lunch  https://aws.amazon.com/blogs/aws/new-amazon-managed-apache-cassandra-service-mcs/

ML works across Tensorflow, PyTorch and mxnet - Sagemaker Studio single pane of glass IDE for machine learning, Sagemaker Notebooks - pairs notebooks with compute, Sagemaker Experiments - Tune, compare, visualize, collect & share models and experiments automatically Sagemaker Debugger - Improve accuracy of models, feature prioritization, metrics for model training, SageMaker Model Monitor - detect concept drift SageMaker AutoML with no loss of visibility or control- CSV(data) -> trains 50 different machine learning models  - with a model leaderboard - notebook with all the models & recipes

Amazon CodeGuru : Auto code reviews + performance profiling - driven by machine learning - input handling, aws best practices, latency & cpu utilization, visualize performance - will find the MOST EXPENSIVE line of code in terms of performance. Installed as an agent on the container . :plus:  web hook for pull requests

Friday, November 1, 2019

Spring RestTemplate Buyer Beware!

TL;DR Be vary of the default RestTemplate injected or manually configured in your existing application. You should leverage HTTP Connection pooling for the RestTemplate which may not be turned  on by default. You  can explicitly configure it with the code sample I provided above. Also the Pool defaults are undersized. change those to a number appropriate to your env.  I set them to max 20 per route. tune per load  also configure connection pool stale connection reaping. Instead of the RestTemplate as the Spring docs advise as of Spring Framework 5.0.

TL;DR based on the multiple enterprise engagements … 

  • The default HTTP client connx. pools must be changed before deploy
  • Don’t forget to set ConnectionRequestTimeout  (defaults to infinity)
  • If possible replace RestTemplate/HttpClient with WebClient, else migrate to okHttpClient which has resolved most observed issues;
  • okHttpClient  has excellent connection pool manager, connection failure and timeouts handling mechanisms..




For instance
    //Wont configure the PoolingHttpClientConnectionManager
    @Bean
    public RestTemplate restTemplate() {
        return new RestTemplate();
    }


    // WILL configure the PoolingHttpClientConnectionManager
    @Bean
    public RestTemplate restTemplate(RestTemplateBuilder builder) {
        return builder.build();
    }

For this to work you need to put HTTPClient or the okHttpclient library on the Classpath


Spring apps leverage the org.springframework.web.client.RestTemplate as a synchronous client to perform HTTP requests. The default configuration of the RestTemplate doesn’t use a connection pool to send requests, it uses a SimpleClientHttpRequestFactory that wraps a standard JDK’s HttpURLConnection opening and closing the connection. This is a problem.  BasicHttpClientConnectionManager can be used for a Low Level, Single Threaded Connection
 https://tech.asimio.net/2016/12/27/Troubleshooting-Spring-RestTemplate-Requests-Timeout.html

Under load Spring RestTemplate client connections are capped at 4 per route. This *blows up under load*. If you see `HTTP status 500` for requests or slow responses please check the HTTPClient configuration and visit the recommendations

## Recommendations
- If you need to have a connection pooling under rest template then you should use different implementation of the ClientHttpRequestFactory that pools the connections. new RestTemplate(new HttpComponentsClientHttpRequestFactory())

- Use the `PoolingHttpClientConnectionManager` to Get and Manage a Pool of Multithreaded Connections. The defaults of the pooling connection manager too small. You should bump UP the MaxTotal, DefaultMaxPerRoute & MaxPerRoute to 20.

- Maximize the utilization of the HTTP Conn Pool
  -  Implement a Custom Keep Alive Strategy
  -  Configure connection evictions to detect idle and expired connections and close them
 -  Read this article https://www.baeldung.com/httpclient-connection-management for connection management.

HttpClientConnectionManager poolingConnManager
  = new PoolingHttpClientConnectionManager();
CloseableHttpClient client
 = HttpClients.custom().setConnectionManager(poolingConnManager)
 .build();

also see https://bitbucket.org/asimio/resttemplate-troubleshooting-svc-2/src/master/src/main/java/com/asimio/api/demo/main/ResttemplateTroubleshootingSvc2Application.java

 As of 5.0, the non-blocking, reactive org.springframework.web.reactive.client.WebClient offers a modern alternative to the RestTemplate with efficient support for both sync and async, as well as streaming scenarios. Always use the *Builder to either create a (or more) RestTemplate or WebClient. Dependencies like spring-cloud-sleuth use the customizer/builder resp.  to add additional features

For greenfield apps pick WebClient over RestTemplate. see

 **The RestTemplate will be deprecated in a future version and will not have major new features added going forward. See the WebClient section of the Spring Framework reference documentation for more details and example code**
 https://www.baeldung.com/spring-5-webclient


## Miscellaneous

  1. For slow requests or for goRouter latency follow https://docs.pivotal.io/pivotalcf/2-5/adminguide/troubleshooting_slow_requests.html and Debugging the Cloud Foundry Routing Tier https://www.youtube.com/watch?v=U5GWgabsxXY
  2. If you encounter a customer that is experiencing an application performance issue (increased latency or decreased throughput or slow requests), try having them run this plugin against the app while it’s under load: https://github.com/cloudfoundry/cpu-entitlement-plugin.
  3. If your Application running on TAS is slow, performing poorly, experiencing high latency and/or decreased throughput then follow debug instructions here  https://community.pivotal.io/s/article/Application-running-on-TAS-is-slow-performing-poorly-experiencing-high-latency-and-or-decreased-throughput


Tuesday, October 29, 2019

Architecture & Services Review Template for 360 degree healthcheck of a Microservice

Do you want to review the health of your system of microservices ? Need a checklist of things to look at as you evaluate the architecture and implementation. Take a look at this all encompassing checklist of things to examine the production readiness and scale of your system of microservices. 


  • Libraries
    • How many unused libraries are there?
    • Are there any libraries that could be replaced by features included with Spring?
  • Connection Pooling
    • How is concurrency handled ?
  • Latency
    • How long does the app take to start up?
    • Is there a meaningful difference in data transmission speed with a high load when using rsockets vs. https?
    • Is there a meaningful difference in data transmission speed when using a reactive tech stack vs. a traditional tech stack?
    • Are there any noticeable areas with inefficient HTTP calls?
    • What is the average response time for the app's network calls?
  • Memory/CPU
    • How much memory does the app use under a high load?. Does it need JVM GC tuning ?
    • How many threads does the app use under a high load?
    • What is the top constraint ? (CPU. Mem, Disk, Network,)
  • Error/Exception Handling
    • How many exceptions does the app usually throw under a high load?
    • What is the mean time between failures?
    • How long does an outage usually last?
  • Code Complexity/Cleanliness
    • What is the highest level of cyclomatic complexity within the app?
    • How many unused classes are in the app?
    • How many unused methods are in the app?
    • Compliance with 15 Factors ?
    • High frequency of code change heat map
    • Sev 1 Production Incidents Review
  • Spring
    • Is there Classpath dependency bloat ?
    • Upgrade to s-boot 2.2 and concomitant dependencies possible ?
  • Resiliency
    • Are circuit breakers and HTTPClients configured correctly
    • Are metrics from Circuit Breakers put in the firehose via micrometer
    • Failure Mode analysis.
  • Observability
    • Are applications logging at the right level
    • Are applications emitting metrics at the right level
    • Is spring-cloud-sleuth enabled for distributed traces ?
    • Configure http healthchecks for the app in Cloud Foundry
  • Performance
    • Is application startup time acceptable. Can this be reduced.
    • Is autoscaling behavior understood in context of downstream dependencies.
    • Policy for autoscaling up and down
  • Higher level Architecture Review

Sunday, October 27, 2019

How do you get Threaddumps and Heapdumps for Java applications running in Cloud Foundry ??

You Cannot.!!  You have hit a classical pain point due to the Java Buildpack using a JRE and not a full JDK .

So the issue is that you cannot cf ssh into the container in PCF and use the jcmd command to trigger a java threaddump. The classical way of resolving high CPU is to take three such threaddumps 30 seconds apart and check to see the threads that are stuck, ones that are not moving or contending on locks or deadlocks etc. You pair this with CPU Profiling information in the VM

NOT able to take a threaddump in PCF is frustrating. WAS/Weblogic had excellent support for getting these artifacts via must-gathers.

So what can you do ? 
You cannot invoke the /threaddump actuator endpoint because that does not provide nearly as much info as a classical threaddump will provide. 

Again this is a problem that anyone who wants to use the JDK tools in an app in PCF faces. Like for instance we want to run the javac command inside the app in PCF. We simply can't due to the above mentioned issue. 

OK So what can be done ... 
A one time custom java buildpack is created rebased on an Open full JDK and not a JRE. This is not sustainable in the long term.  You will need to restage the app with this custom Full JDK Java buildpack. 
- The JDK tooling (jcmd, jmap and other command line tools) have to be trojan horsed into the app via a side-car container or something like a pcfshell https://github.com/tfynes-pivotal/pcfshell or the app has to carry the executable with it. 
- Another option is that app itself carries a /threaddump endpoint via a spring boot actuator although if the app is dying due to OOM or high CPU this seldom works
- If the app is crashing due to an OOM it writes out a histogram and a cause of failure. In such a case enable verbose GC logging to stdout so that you can collect and visualize the GC logs and 2. you can configure a persistent volume bind for the Java buildpack to write the core file to a persistent volume oom-killer  jre-docs
Existence of a single bound Volume Service will result in Terminal heap dumps being written.
- Use flame graphs in PCF to debug high CPU. This requires some investigation. 

Thursday, October 24, 2019

Migrate away from IBM Integration Bus

Monolith  ---------------------> DB

Step 1.
Monolith  ----------> DB |  ACL |  Microservice1     ------> new DB (Read only)
 - all data is migrated in read-through from old DB to new DB via ACL
 - Migrate 90% of the data like this
 - newDB <-- sync --> oldDB [needs Synchronization]

---

Migrate off of IIB
 - Rapid migration off of IIB
 - Take the custom code > wrap it in s-boot and don't change the data model
 - No local database .. all IIB converted apps talking to data model
 - [x] Conformist pattern ... then no ACL
 - [y] Evolve the API and add new consumers then add ACL and use your domain model. no database.

 (1) - Proxy off of IIB <READ>
 (2) - Selective re-examine data strategy based on the app <READ|WRITE>

---

DDD is BROKEN!

The theory of domain driven design created by Eric Evans in his seminal book  Domain Driven Design (DDD) was published in August 30, 2003. Now DDD is such a dense tome that it requires an average senior software engineer two tries to read after which you wonder how exactly you apply this to running software. Thereafter you start reading the Vaughn Vernon's red book - Implementing DDD to figure out the implementation of these patterns in code. At this point most software engineers are still struggling to apply the tactical and strategic patterns of domain driven design. This was the state of DDD around 2010.

Enter Adrian Cockcroft of Netflix fame who along with Martin Fowler sparked the microservices revolution and a renewed interest in DDD as the theoretical underpinning for microservices. This jazzed up everyone, since no one had a clue about how to structure the boundaries and domain for individual microservices. What is the correct way to design the boundaries of your services ? Bounded Contexts and sub-domains and context maps from DDD came to the rescue. and provided a basis to structure your system of microservices.

So we have the theory to now split and structure microservices. This again is NOT Enough. A lot of architects floundered in trying to figure out how to transform from a massive monolith to an event driven choreographed system of microservices. They struggled defining the subdomains and the bounded contexts.

Enter Alberto Brandolini who figured out that Event Storming is the only way that merges the people and technical aspects, the tactical and strategic aspects to visualize domains. Event Storming is a cross functional facilitation technique for revealing the bounded contexts, microservices, vertical Slices, trouble spots and starting points for a system or business process. Event Storming and other techniques that I mention later allowed architects and product owners to practice strategic DDD. There was no systematic way to practice strategic DDD to software before event storming. Alberto's influence on DDD in seminal as one who democratized it and made it available to masses. There is also an analogy to legos here. In the 1950's legos were introduced in Scandinavia, however their sales were struggling. Lego was far rom the powerhouse it was today.  It is only when the Lego company started shipping instructions to build the lego sets, that sales took off. This is key  - the packaging, the instructions that allowed a seven year old to build the set on their own and get a sense of accomplishment without nagging his/her parents.

Now having practiced event storming for a couple of years , we realized that incremental notation is key. If you get stuck on creating a color coordinated combination of aggregates, commands and events , read models, UI, data and policies it can pigeon hole and restrict the domain model and lead to a tunnel effect. Domain Events are front and center for event storming, everything else is secondary and needs to be added incrementally.  The gap from event storming a system to an actual backlog of stories is HUGE. If you follow the textbook definitions of event storming you end up an event sourced CQRS system which most developers struggle to implement and maintain. A mistake i have lived with. There are multiple forms of event driven architecture and choreography what-is-event-driven. We want to start with the easier incantations and then graduate to the top level of a full blown event sourced system.

This led to the  SWIFT method that leverages a technique called Boris that uses graph theory to model the relationships between the capabilities in a system. This process generates information about how the system "wants to be designed" and attempts to avoid pitfalls such as premature solutioning. At the end of a Boris Exercise, Services, APIs, Data and Event Choreography and a backlog of work starts becoming obvious.

After Boris it is critical to run multiple modeling exercises and then determine MVPs for the vertical flows. In order to determine the right MVPs for your system you have to consider thin vertical end to end slices where these domains interact with one another. You have to prioritize the thin slices based on technical effort, risk and business value. The slices encompass a sub  section of events . The MVPs map a path from strangling the monolith and leveraging tactical patterns to interact with the new domains and services. There are multiple techniques that can be applied here including domain story telling and user story impact mapping.

Vertical slices are identified by choosing short, domain event flows in the core domain and defining the architectural components required to produce those events. Slice by slice, translate the domain model into microservices that use APIs, message queues, etc., that will run on the platform. Finally, a set of user stories is defined and mapped to releases or MVP.

After performing the mapping of user stories that realize the tactical patterns to MVPs, we now have a concrete backlog that developers can start with and iterate on.

Here is the full sequence of steps that help with decomposing a monolith:
1. Define Objectives and Key Results (OKR) for the app modernization effort.
2. Event storm the application and identify bounded contexts.
3. Pick several short domain event flows in the core domain that constitute a vertical slice.
4. Create Boris diagrams that define the relationships between the domains for an end to end slice
5. Perform SNAP analysis to score the effort and define data, API and messaging interfaces
6. Create a backlog of prioritize user stories tied back to OKR.
7. Impact Map user stories to MVP or releases.

At this point the benefits and promises of DDD become real and the theory of DDD laid out by Eric Evans becomes a pragmatic living software system.

Happy Modeling.
Rohit Kelapure