Rajnish Noonia

Category: Architecture

  • Enterprise configuration management

    Almost every application requires some form of configuration information. This information can be as simple as a database connection string or as complex as multipart and hierarchical user preference information. How and where to store an application’s configuration data are questions you often face as a developer.

    Any large enterprise application has many moving blocks. They all need to be configured for a proper working of the application. As the application size increases or for scalability the same configuration has to be repeated in different applications. For most applications once the configuration has been changed the application needs to be restarted.

    Sample Code (create a blank console project, add json.net nuget)

    using Newtonsoft.Json;
    using Newtonsoft.Json.Linq;
    using System.Collections.Generic;
    using System.Linq;
    
    namespace CM
    {
        // configuration management API, 
        //1. allow clients specify typesafe models for configuations
        //2. Store flat data on server (table) which is easy to edit
        //3. can be extended to have inheritance of values (overrides)
        //4. can be extended to lock / unlock certain property by admins etc.
        
        class Program
        {
            //Sample configuration model
            internal class SampleConfigModal
            {
                public SampleConfigModal()
                {
                    Address = new Address();
                }
                public string Name { get; set; }
                public Address Address { get; set; }
                public int Age { get; set; }
    
    
            }
            public class Address
            {
                public string Street { get; set; }
    
            }
    
            // this is how client api will look like
            static void Main(string[] args)
            {
                // sample client code
                var data = new SampleConfigModal() { Name = "Rajnish", Age = 18, Address = new Address() { Street = "Oxley" } };
    
                // save Configuration
                SaveConfiguration("app", "section", data);
    
                //get Configuration
                var data2 = GetConfiguration("app", "section");
            }
    
            // Client side framework api -> call to rest end point
            private static T GetConfiguration(string appName, string SectionName) where T : new()
            {
                var defaultValue = new T();
                var samplePayload = JsonConvert.SerializeObject(defaultValue);
                var payload = GetConfiguration(appName, SectionName, samplePayload);
                return JsonConvert.DeserializeObject(payload);
            }
    
            // Client side framework api -> call to rest end point
            private static void SaveConfiguration(string appName, string SectionName, T data)
            {
                var payload = Newtonsoft.Json.JsonConvert.SerializeObject(data);
                SaveConfigurationa(appName, SectionName, payload);
            }
    
    
            //---------------------------------------- Server Code -------------------------- 
            //-------------- server has no knowledge of configuration structure or model
    
            private static Dictionary<string, string> storage;
    
            private static void SaveConfigurationa(string appName, string SectionName, string payload, string enumForHierarchyLevel = null)
            {
                // transformer
                var section = string.Format("{0}.{1}", appName, SectionName);
                var data = (JObject)JsonConvert.DeserializeObject(payload);
                var keyValueData = Flatten(data, section);
    
                // check if user has permission for level overrides
                // store with proper overides
    
                // store the flat list in sql or data 
                //| KEY |           |Value|            |OverrideType| - default,sysadmin,appadmin,groups,user etc
                //app.section.Name, Rajnish
                //app.section.Address.Street, Oxley
                //app.section.Age, 18
    
                storage = keyValueData;
            }
    
            private static string GetConfiguration(string appName, string SectionName, string samplePayload)
            {
                var section = string.Format("{0}.{1}", appName, SectionName);
                var data = (JObject)JsonConvert.DeserializeObject(samplePayload);
                var keyValueSample = Flatten(data, section);
                // update data from sql or data store
                // apply property override rules and get value from overrides if exists
                var keyValueData = keyValueSample.Select(x => new KeyValuePair<string, string>(x.Key, storage[x.Key]));
    
                //read these
                //app.section.Name, Rajnish
                //app.section.Address.Street, Oxley
                //app.section.Age, 18
    
                UnFlatten(data, section, keyValueData);
    
                var formatedData = JsonConvert.SerializeObject(data);
                /*
                 * {
                      "Name": "Rajnish",
                      "Address": {
                        "Street": "Oxley"
                      },
                      "Age": "18"
                    }
                 * */
                return formatedData;
    
            }
            
            // Server side json helper
    
            private static void UnFlatten(JObject jsonObject, string prefix, IEnumerable<KeyValuePair<string, string>> data)
            {
                foreach (var item in data)
                {
                    var keyName = item.Key.Substring(prefix.Length + 1);
                    var storageValue = item.Value;
                    if (keyName.Contains("."))
                    {
                        var keys = keyName.Split('.');
                        var jtoken = (JToken)jsonObject;
                        foreach (var k in keys)
                        {
                            jtoken = jtoken.SelectToken(k);
                        }
                        ((JValue)jtoken).Value = storageValue;
                    }
                    else
                    {
                        jsonObject[keyName] = storageValue;
                    }
                }
            }
    
            private static Dictionary<string, string> Flatten(JObject jsonObject, string prefix)
            {
    
                IEnumerable jTokens = jsonObject.Descendants().Where(p => p.Count() == 0);
                Dictionary<string, string> results = jTokens.Aggregate(new Dictionary<string, string>(), (properties, jToken) =>
                {
                    properties.Add(string.Format("{0}.{1}", prefix, jToken.Path), jToken.ToString());
                    return properties;
                });
                return results;
            }
        }
    }
    
    
  • BoundedContext – DDD

    Earlier in the article Software Architecture Patterns we briefly discussed domain driven design. In this article we will take a real example and dive into best practices and design of solution based on hypothetical problem or scenario.

    Bounded Context is a central pattern in Domain-Driven Design and It is the focus of DDD’s strategic design section which is all about dealing with large models and teams. The DDD deals with large models by dividing them into different Bounded Contexts and being explicit about their interrelationships.

    Strategic design  deals with situations that arise in complex systems, larger organizations, interactions with external system.Strategic design decisions are made by teams, or even between teams. Strategic design enables the goals of DDD to be realized on a larger scale, for a big system or in an application that fits in an enterprise-wide network.

    DDD is about designing software based on models of the underlying domain. A model acts as a Ubiquitous language to help communication between software developers and domain experts. It also acts as the conceptual foundation for the design of the software itself.

    It is hard to model a larger domain and build a single unified model. In real world, small domain models are build and together they represent the larger domain. Now lets image a real world example from electricity utility – smart meters ! –  here the word “meter” meant subtly different things to different domain experts coming from different parts of the organization. Lets try to understand the domain and try to break into sub domain models.

    Smart meters are the next generation of gas and electricity meters and offer a range of intelligent functions.The smart metering system is made up of: one electricity smart meter, one gas smart meter, a communications hub and an in-home display unit-the smart energy monitor on which you can view your energy.Smart meters measure actual, total gas and electricity usage and put consumers in control of their energy use, allowing them to adopt energy efficiency measures that can help save money on their energy bills.

     

    smart meter

     

    Now lets split the into domains and sub domains

    smart domain

     

    In the above diagram the subdomain build on foundation however one sub domain is interrelated to one or more other sub domain. They don’t exists in isolation in real world. The total unification of the domain model for a large system will not be feasible. So instead DDD divides up a large system into Bounded Contexts, each of which can have a unified model.

    A bounded context typically represents a slice of the overall system with clearly defined boundaries separating it from other bounded contexts within the system. If a bounded context is implemented by following the DDD approach, the bounded context will have its own domain model and its own ubiquitous language.

    A bounded context is the context for one particular domain model. Similarly, each bounded context (if implemented following the DDD approach) has its own ubiquitous language, or at least its own dialect of the domain’s ubiquitous language, entities, services etc. as shown below.

    Bounded Context

    Bounded Contexts have both unrelated concepts – such as a support ticket only existing in a customer support context, but also share concepts such as products and customers both exists in sales and support contexts.

    A large complex system can have multiple bounded contexts that interact with one another in various ways. A context map is the documentation that describes the relationships between these bounded contexts. It might be in the form of diagrams, tables, or text.

    In the next series we will dive into how we can apply CRQS (Command Query Responsibility Segregation Pattern) and vertically slice the layered architecture to deliver the highly scalable yet composite solution.

  • Software Architecture Patterns

    Extreme Programming (XP) is one of the more well known Agile methodologies. It is a programmer-centric methodology that emphasizes technical practices to promote skillful development through frequent delivery of working software.This methodology takes “best practices” to extreme levels and that’s why its named as Extreme Programming. Code reviews are a good example of Extreme programming. If code reviews are good, then doing constant code reviews would be extreme; but would it be better? This led to practices such as pair-programming and refactoring, which encourage the development of simple, effective designs, oriented in a way that optimizes business value.

    Extreme Programming defines 4 basic activities (coding, testing, listening & designing) and several practices like Pair Programming, Planning game, Test driven development, continuous integration, design improvement, coding standards, collective code, simple design etc.

    Projects suited to Extreme Programming are those that:

    • Involve new or prototype technology, where the requirements change rapidly, or some development is required to discover unforeseen implementation problems
    • Are research projects, where the resulting work is not the software product itself, but domain knowledge
    • Are small and more easily managed through informal methods

    Below are some software development process based on a concept of Extreme Programming and are about how to approach your design.

    Test driven design (TDD)

    tdd

    TDD is a software development process that relies on the repetition of a very short development cycle: requirements are turned into very specific test cases, then the software is improved to pass the new tests, only. It offers them a technique to explore the concepts behind the customers requirements, questioning that requirement and uncovering likely pitfalls. The developer can deliver these benefits without spending valuable time building and perfecting a graphical user interface. it stops developers from over engineer the product and encourage them to think from different prospective.

    TDD relies on the repetition of a very short development cycle :

    • Write an automated test case that defines a new feature – no code yet, so test will fail
    • Produce the minimum amount of code to pass that test
    • Refactor the new code to acceptable standards.

    Domain driven design (DDD)

    DDD is the process of being informed about the Domain before each cycle of touching code. Domain is a set of functionality that you are attempting to mimic that lies outside of your application.Domain Driven Design (DDD) is about mapping business domain concepts into software artifacts.Driven Design (DDD) focuses on the core model (the domain) and tries to keep other stuff like UI’s and databases separate.Domain Driven Design is all about understanding the customer real business need and emphases focuses more into the business need not focusing on the technology.

    It promotes important agile principles:-

    • Maintain the projects primary focus on the core domain of the delivery
    • Use models to refine a complex design and
    • Get the key team members together to collaborate deeply to derive their designs.

    Domain modeling and DDD play a important role in Enterprise Architecture (EA). Since one of the goals of EA is to align IT with the business units, the domain model which is the representation of business entities, becomes a core part of EA. This is why most of the EA components (business or infrastructural) should be designed and implemented around the domain model. Domain driven design is a key element of Service Oriented Architecture (SOA) because it helps in encapsulating the business logic and rules in domain objects. The domain model also provides the language and context with which the service contract can be defined.

    Domain driven design effort begins where domain modeling ends.

    There should be more focus on domain objects than services in the domain model.

    • Start with domain entities and domain logic.
    • Start without a service layer initially and only add services where the logic doesn’t belong in any domain entity or value object.
    • Use Ubiquitous Language, Design by Contract (DbC), Automated Tests, CI and Refactoring to make the implementation as closely aligned as possible with the domain model.

    From the design and implementation stand-point, a typical DDD framework should support the following features.

    • It should be a POCO based framework.
    • It should support the design and implementation of a business domain model using the DDD concepts.

    Bounded Context is a central pattern in Domain-Driven Design and It is the focus of DDD’s strategic design section which is all about dealing with large models and teams. – for more details on bounded context continue reading ddd here – series 2 of DDD.

    Behaviour driven design (BDD)

    BDD is a software development process based on Test-driven Development (TDD), that combines the general techniques and principles of TDD with ideas from Domain-driven Design (DDD) and Object-oriented Analysis and Design to provide software developers and business analysts with shared tools and a shared process to collaborate on software development, with the aim of delivering “software that matters”.While it is a refinement to TDD, it concentrates in understanding the user’s behaviour, and yields nicely to a good acceptance of the end system. In the though process means thinking from outside the system in. The benefit is that it offers a more precise and organized conversation between developers and domain experts

    BDD is also often heralded because BDD testing tools can be arguably more human readable to non-developers such as Domain Experts

    Event driven architecture (EDA)

    -TODO

    Command Query Responsibility Segregation Pattern (CQRS)

    -TODO

  • Cloud Computing

     

    Cloud

    You’re probably using cloud computing right now, even if you don’t realize it. If you use an online service to send emails, edit documents, watch films or TV, listen to music, play games, or store pictures and other files, it’s likely that cloud computing is making it all possible behind the scenes.

    Cloud computing stack

    Most cloud computing services fall into four broad categories: On Premises, infrastructure as a service (IaaS), platform as a service (PaaS) and software as a service (SaaS).

    Cloud Solutions Model

    IaaS  – Infrastructure as service

    This is where pre-configured hardware is provided via a virtualised interface or hypervisor. There is no high level infrastructure software provided such as an operating system, this must be provided by the buyer embedded with their own virtual applications.

    PaaS – Platform as service

    PaaS goes a stage further and includes the operating environment included the operating system and application services. PaaS suits organisations that are committed to a given development environment for a given application but like the idea of someone else maintaining the deployment platform for them.

    SaaS – Software as service

    Saas offers fully functional applications on-demand to provide specific services such as email management, CRM, web conferencing and an increasingly wide range of other applications & services.

    Type of Cloud deployment

    Based on the security and management required, the clouds can be built in following three ways to suit the needs of the businesses:

    Public cloud
    Public clouds are owned and operated by a third-party cloud service provider, which delivers computing resources such as servers and storage over the Internet. Microsoft Azure is an example of a public cloud. With a public cloud, all hardware, software and other supporting infrastructure are owned and managed by the cloud provider. You access these services and manage your account using a web browser.

    Private cloud
    A private cloud refers to cloud computing resources used exclusively by a single business or organisation. A private cloud can be physically located on the company’s on-site data centre. Some companies also pay third-party service providers to host their private cloud. A private cloud is one in which the services and infrastructure are maintained on a private network.

    Hybrid cloud
    Hybrid clouds combine public and private clouds, bound together by technology that allows data and applications to be shared between them. By allowing data and applications to move between private and public clouds, hybrid cloud gives businesses greater flexibility and more deployment options.

    Community Cloud

    Type of cloud hosting in which the setup is mutually shared between many organisations that belong to a particular community, i.e. banks and trading firms. It is a multi-tenant setup that is shared among several organisations that belong to a specific group which has similar computing apprehensions. The community members generally share similar privacy, performance and security concerns.

  • TPL Dataflow – Concurrent Programming

    TPL DataFlow

    TPL Dataflow is an in-process actor library on top of the Task Parallel Library enabling more robust concurrent programming.

    Parallel computing is a form of computation in which multiple operations are carried out simultaneously.Parallel computing is closely related to asynchronous programming, using many of the same core concepts and support. Asynchronous programming is an approach to writing code that involves invoking operations such that they don’t block the current thread of execution.Many personal computers and workstations have two or four or 8 cores (that is, CPUs) that enable multiple threads to be executed simultaneously. Computers in the near future are expected to have significantly more cores. To take advantage of the hardware of today and tomorrow, you can parallelize your code to distribute work across multiple processors. In the past, parallelization required low-level manipulation of threads and locks.

    The purpose of the TPL is to make developers more productive by simplifying the process of adding parallelism and concurrency to applications. The TPL scales the degree of concurrency dynamically to most efficiently use all the processors that are available. In addition, the TPL handles the partitioning of the work, the scheduling of threads on the ThreadPool, cancellation support, state management, and other low-level details. By using TPL, you can maximize the performance of your code while focusing on the work that your program is designed to accomplish.

    Data parallelism refers to scenarios in which the same operation is performed concurrently (that is, in parallel) on elements in a source collection or array. In data parallel operations, the source collection is partitioned so that multiple threads can operate on different segments concurrently.

    The Task Parallel Library (TPL) is based on the concept of a task, which represents an asynchronous operation. In some ways, a task resembles a thread or ThreadPool work item, but at a higher level of abstraction. The term task parallelism refers to one or more independent tasks running concurrently. Tasks provide two primary benefits:More efficient and more scalable use of system resources & More programmatic control than is possible with a thread or work item.

    The concurrency models we has discussed so far have the notion of shared state (data) in common.Shared state can be accessed by multiple threads at the same time and must be thus protected, either by locking or by using transactions. Both, mutability and sharing of state are not just inherent for these models, they are also inherent for the complexities.Unfortunately, programmers have found it very difficult to reliably build robust multi-threaded applications using the shared data and locks model, especially as applications grow in size and complexity.Making things worse, testing is not reliable with multi-threaded code. Since threads are non-deterministic, you might successfully test a program one thousand times, yet still the program could go wrong the first time it runs on a customer’s machine.

    We now have a look at an entirely different approach that bans the notion of shared state altogether. State is still mutable, however it is exclusively coupled to single entities that are allowed to alter it, so-called actors.The actor model in computer science is a mathematical model of concurrent computation that treats “actors” as the universal primitives of concurrent digital computation: in response to a message that it receives, an actor can make local decisions, create more actors, send more messages, and determine how to respond to the next message received.For communication, the actor model uses asynchronous message passing. In particular, it does not use any intermediate entities such as channels. Instead, each actor possesses a mailbox and can be addressed. These addresses are not to be confused with identities, and each actor can have no, one or multiple addresses. When an actor sends a message, it must know the address of the recipient. In addition, actors are allowed to send messages to themselves, which they will receive and handle later in a future step.

    The Task Parallel Library (TPL) provides dataflow components to help increase the robustness of concurrency-enabled applications. These dataflow components are collectively referred to as the TPL Dataflow Library. This dataflow model promotes actor-based programming by providing in-process message passing for coarse-grained dataflow and pipelining tasks.The TPL Dataflow Library provides a foundation for message passing and parallelizing CPU-intensive and I/O-intensive applications that have high throughput and low latency. It also gives you explicit control over how data is buffered and moves around the system.

    If you want to scale your application beyond single machine or process then ServiceBus (NServiceBus, Microsoft Azure, etc) are the best candidates. These are designed around message oriented architecture and you can achieve highly reliable, available and scalable application. As an architect i always focus on reliability, after all, a highly available and scalable service that produces unreliable results isn’t very valuable. – We will need another post to cover the in-depth of service bus..

    TPL Dataflow (TDF) is a library for building concurrent applications. It promotes actor/agent-oriented designs through primitives for in-process message passing, dataflow, and pipelining. I have been playing with dataflow since its CTP was released and i found its very use in cases where you have to process data in form of a pipeline.With just few in-build blocks you can easily and quickly build concurrent app..

    The primitive blocks provided by dataflow are

    • Buffering Blocks – Holds data for use by data consumers.
      • BufferBock(T) – FIFO queue of message that can be written to multiple sources or read from by multiple targets.
      • BroadcastBlock(T) – Used when you must pass multiple messages to another component.
      • WriteOnceBlock(T) – similar to broadcastblock except object can be written to one time only
    • Execution Block – call a user provided delegate for each piece of received data
      • ActionBlock(t) – calls a delegate when it receives a data – excepts synchronous or asynchronous delegates
      • TransformBlock(Tinout,TOutput) – call function delegates to transform the incoming message to another type- excepts synchronous or asynchronous delegates
      • TransformManyBlock(TInput , TOutput) – similar to TransformBlock except it can produce zero or more output values for each input value, instead of only one output value for each input value. – excepts synchronous or asynchronous delegates

    Degree of Parallelism

    Every ActionBlock<TInput>, TransformBlock<TInput, TOutput>, and TransformManyBlock<TInput, TOutput> object buffers input messages until the block is ready to process them. By default, these classes process messages in the order in which they are received, one message at a time. You can also specify the degree of parallelism to enable ActionBlock<TInput>, TransformBlock<TInput, TOutput> and TransformManyBlock<TInput, TOutput> objects to process multiple messages concurrently.

    Now lets look at the implementation details of a web crawler

    Request Buffer
    |
    PageDownload
    / Save     ParseLink
    |
    RaiseLinkFound

    The messages to download a Url is received in the request buffer which is downloaded by a TranformBlock and converted into type safe page type message. The page message is then broadcasted using Broadcast block, which is further received by save ActionBlock and ParseLink Block. The parse link block parses the urls in the page and if they belong to same page it will raise an event for each url. The consumer of engine will receive the url and if its a new URL it will be posted back to engine.. The save action block will save the page to disk.

    This is very basic example but you can see the message based approach is much more simpler than a shared resource + threading approach.

    The TPL dataflow is good in case you don’t want to scale the solution beyond single machine as it offer a in process message base approach.With a proper service bus like NServiceBus you can scale out the solution beyond single machine and multiple servers could process the request to achieve the high throughput and off-course with easy to build,maintain clean code base.

    In production there are more things you have to take care like logging, error handling, transactions, unexpected failure recovery, dependency injection,loosely coupled components, extensibility, scalability and so on.. Frameworks like NServiceBus provides all these features alone with API to handle more complex business problems.

    Download Code : Here (Partially finished but working POC)