What can Jenkins do for you?

Topic: “What can Jenkins do for you?” might sound a bit old fashioned and cliched as Jenkins has been around for a while but it has very varied capabilities via plugins & build pipelines to manage many things. Brief list of capabilities which in no way are exhaustive are given below:

  1. Continuous build management
  2. Continuous deployment
  3. Continuous testing
  4. Continuous quality checks and code scans
  5. Continuous security testing
  6. Continous license checks
  7. Continous Kubernetes, cloud & docker deployment / monitoring
  8. Continuous email notifications for events
  9. Integration with JIRA
  10. Integration with notification systems
  11. Continuous monitoring
  12. Continuous reports & test results analysis

Key concepts, documentation & keywords in Kafka – Part 1

Here are some important concepts, documentation and keywords of Kafka that you can refer and learn. There are two major flavors of Kafka – Apache Kafka & Confluent Kafka, I have listed major keywords, documentation and concepts from both here:

  • Broker
  • Zookeeper
  • kSQL
  • REST-Proxy
  • Schema-Registry
  • Connectors
  • Operator
  • Control Center
  • Streams
  • Topics
  • Consumers
  • Producers
  • Partitions
  • Offset
  • Log
  • Node
  • Replica
  • Message
  • Leader
  • Follower
  • Replicator
  • Schema management
  • Confluent Hub
  • Events
  • Associated keywords in today’s cloud deployments: Docker containers, Kubernetes, Ansible, Security

Associated documentation:

Building data models that everyone can understand and more importantly believe

Building data models that everyone can understand and more importantly believe. Faculty Article – Author: Mr. Balakrishnan Unny & Mr. Neil Harwani. Thank you Sapience – IMNU’s (Nirma University) Alumni Newsletter for publishing our article in Changing Times 2.0 (A Special Edition).

Productivity hacks for Architects / Designers / Tech Leads

As per my experience, the biggest productivity hacks for Architects / Designers / Tech Leads are not to decide the variables / class names / loops / scope / data types / exception handling / object relational mapping & so on – they definitely are important and should be done, but so are the below points:

1. Design patterns

2. What is the code for?

3. Functional to technical mapping

4. Solution creation

5. Pseudocode & logic steps

6. Logic of solution for design / programming problems

7. Co-ordination with stakeholders & communication

8. Code review

9. Logic review of programmed modules

10. Architecture / Design thoughts

11. Knowledge updation around tools / products / frameworks usage

12. Time management of developers

13. Task management of developers

14. Solving problems in design

15. Programming standards management

16. Technical best practices management

17. New technology exploration

18. Helping sales, presales & practice

19. Working on POCs, solutions, products and accelerators

20. Updating oneself with the current happening in industry and domain

21. Automation, Security, Testing, Deployment, Continuous integration / deployment, Integrations, Logging, User Interface / User Experience, Application monitoring, Support structure, Clustering / Auto-scaling, Non functional requirements and other such important areas

22. Establish collaboration / teamwork among technical staff working with them

23. Right documentation and knowledge sharing practices

Many get stuck in only programming, that is definitely something we all love and do, but you should be dividing your time as an Architect / Tech Lead / Designer between programming and above tasks equally. Current enterprise softwares are complex and you can’t achieve much without collaboration and above form an important link for productivity in complex, large team projects.

#architecture #design #technicallead #solutionsarchitect

Email me: Neil@TechAndTrain.com

Visit my creations:

  • www.TechAndTrain.com
  • www.QandA.in
  • www.TechTower.in

Data Analysis Process in Analytics / Data Science

This article is based on understanding from Wikipedia article on Data Analysis & my experiences in Data Science / Analytics / AI / ML – https://en.wikipedia.org/wiki/Data_analysis

Various areas like Data Mining, Predictive Analysis, Exploratory Data Analysis, Text Analytics, Business Intelligence, Confirmatory Data Analysis and Data Visualization overlap with this area

Before starting your journey on solving an industry or academic or research problem in Data Science / Analytics / AI / ML / Decision Science, a fundamental step where many students & professionals struggle is Data Analysis. In this article, I am providing a step by step approach on analyzing your data. Directly starting with programming of various algorithms or neural network on your data could at times be counterproductive and should be avoided. Initial stage should involve robust data analysis via steps given below followed by model building which can include custom or already proven algorithms or a derivative of some popular models. For each of the points discussed below, I have added additional information on top of interpretation of Wikipedia information based on my experience in industry towards the end of each of the points or I have added new points post the interpretations.

Your steps for data analysis should generally be:

  1. Setup your data analysis process at a high level with your objectives – inspecting data, cleaning it, processing (could include dimensionality reduction / feature engineering), transformation, modelling and communicating it. Many forget the functional and feedback loop in this process setup to improve data quality – that must be included too.
  2. Next step is in understanding the data in terms of what is it telling us. Data could be quantitative style numbers or textual or a mix of it. Treatment for all three is different. For quantitative / numerical data, we try to understand whether it is time-series, ranking, part to whole, deviation, frequency distribution, correlation, nominal or geographical or geospatial data. For textual or mixed type of data we need to use approaches of text mining, sentiment analysis, natural language processing to get insights around frequency of words, influential words & sentences by weight, trends, categories, clusters and more. Most of this article revolves around quantitative or numerical data perse and not textual data. I have provided a very brief idea on textual data analysis here in this point.
  3. Next step would be to have the quantitative techniques being applied on the data in terms of sanity, audit / reconciliation of totals via formulas, relationships between data, checking things like whether variables in data are related in terms of correlation / sufficiency / necessity / etc. I would suggest using R Studio or similar tool for this step.
  4. Post this we want to actually perform actions like filtering, sorting, checking range and classes, summary, clusters, relationships, context, extremes, etc. At this stage, exploratory data analysis techniques come in very handy where we use various libraries which provide graphical representation. Excel & Tableau come in handy here.
  5. Our next step will be to check for biases, deciphering facts & opinions, deciphering any numerical incorrect / irrelevant inferences which are being projected and need correction / improvement. This needs detailed study of data from domain / functional perspective and applying statistical analysis on it. Working with a business / functional consultant in this phase is especially useful.
  6. Some areas which we need to take care of include quality of data, quality of measurements, transformations various variables / observations into log scale or others like what we have on richter scale for earthquakes, mapping to objectives and characteristics. This is an intuitive step where visualizing data through various transformations in R / Python / etc. using libraries like Ggplot2, Plotly, Matplotlib, etc. helps.
  7. Next comes checking outliers, missing values, randomness, analysis & plotting various of charts based on type of data whether categorical or continuous. This is statistical analysis & visualization where I find R to be most suited.
  8. Building models around our data analysis steps could involve linear, non-linear models and checking values via hypothesis testing and mapping to algorithms to process, predict, cluster, find trends and so on. Products / tools like R / Python with libraries like Scikit learn, Numpy, Pandas, MLR, Caret, Keras, TensorFlow, etc. help here
  9. While running the models take care of cross-validation of data & sensitivity analysis – This can generally be done using some options in model training & testing phase for supervised learning.
  10. Feedback loop to circle and improve data & results, accuracy analysis and improvement, pipeline building, interpretation of results & functional mapping to domain are additional things that we need to consider on top of the basics given in Wikipedia article. Also, things like dimensionality reduction techniques like PCA, SVD and such need to be explored in detail as they are helpful in this analysis.

Additional information on top of what is in Wikipedia article:

  1. Explainable AI / ML – https://en.wikipedia.org/wiki/Explainable_artificial_intelligence
  2. Interpretable ML – https://statmodeling.stat.columbia.edu/2018/10/30/explainable-ml-versus-interpretable-ml/
  3. Tools / languages / products to use: R, Python, Pandas, Numpy, Tableau and so on
  4. EDA – https://en.wikipedia.org/wiki/Exploratory_data_analysis
  5. Which chart to use – https://www.tableau.com/learn/whitepapers/which-chart-or-graph-is-right-for-you
  6. List of charts – https://python-graph-gallery.com/all-charts/
  7. Confirmatory data analysis – https://en.wikipedia.org/wiki/Statistical_hypothesis_testing
  8. Singular Value Decomposition – https://en.wikipedia.org/wiki/Singular_value_decomposition
  9. Dimensionality Reduction – https://en.wikipedia.org/wiki/Dimensionality_reduction

Email me: Neil@TechAndTrain.com

Visit my creations:

  • www.TechAndTrain.com
  • www.QandA.in
  • www.TechTower.in

What are we doing in AI / ML / Data Science / Decision Science / Analytics World? – Glossary

Over the last few years I have explored, programmed, worked in, researched and taught Data Science / AI / ML / Analytics / Decision Science to multiple students and with many software professionals. I have collected many keywords that you can google and explore. This will help you to keep pace and learn about things happening is these areas. It’s like a glossary of words to search over internet. It’s a mix and match of technologies, algorithms, concepts, AI / ML / Information Technology terms, BigData words and so on in no particular order. I will keep expanding this till it’s a relatively exhaustive list.

  • Automatic Machine Learning
  • Transfer Learning
  • Explainable Machine Learning
  • Keras
  • PyTorch
  • MLR
  • R
  • Python
  • Ggplot2
  • MathplotLib
  • MLib
  • Spark
  • Hadoop
  • Tableau
  • Chatbots
  • Talend
  • MongoDB
  • Neo4j
  • Kafka
  • ELK
  • NoSQL
  • Cassandra
  • AWS SageMaker
  • SVM
  • Decision Trees
  • Regression: Logistic, Multiple, Simple Linear, Polynomial
  • Scikit Learn
  • KNIME
  • BERT
  • NLG
  • NLP
  • Random Forest
  • Hyper parameters
  • Boosting
  • Association rules / mining – Apriori, FP-Growth
  • Data mining
  • OpenCV
  • Self driving cars
  • AI / Memory embedded SOCs, GPUs, TPUs
  • Neural engine chipsets
  • Neural Networks
  • Deep Learning
  • EDA
  • Statistical & Algorithmic modelling
  • Sampling
  • Probability distributions
  • Hypothesis testing
  • Intervals, extrapolation, interpolation
  • Scaling
  • Normalization
  • Agents, search, constraint satisfaction
  • Rules based systems
  • Semantic net
  • Propositional logic
  • Fuzzy reasoning
  • Probabilistic learning
  • First order logic
  • Game theory
  • Pipeline building
  • Ludwig
  • Bayesian belief networks
  • Anaconda Navigator
  • Jupyter
  • Synthetic data
  • Google dataset search
  • Kaggle
  • CNN / RNN / Feed forward / Back propagation / Multi-layer
  • Tensorflow
  • Deepfakes
  • KNN
  • K means clustering
  • Naive Bayes
  • Dimensionality reduction
  • Feature engineering
  • Supervised, unsupervised & reinforcement learning
  • Markov model
  • Time series
  • Categorical & Continuous data
  • Imputation
  • Data analysis
  • Classification / Clustering / Trees / Hyperplane
  • Differential calculus
  • Testing & training data
  • Visualization
  • Missing data treatment
  • Scipy
  • Pandas
  • LightGBM
  • Numpy
  • Dplyr
  • Google Collaboratory
  • PyCharm
  • Plotly
  • Shiny
  • Caret
  • NLTK, Stanford NLP, OpenNLP
  • Artificial intelligence
  • SQL / PLSQL
  • Data warehousing
  • Cognitive computing
  • Coral
  • Arduino
  • Raspberry Pi
  • RTOS
  • DARPA Spectrum Challenge
  • 100 page ML book
  • Equations, Functions, and Graphs
  • Differentiation and Optimization
  • Vectors and Matrices
  • Statistics and Probability
  • Operations management & research
  • Unstructured, semi-structured & structured data
  • Five Vs
  • Descriptive, Predictive & Prescriptive analytics
  • Model accuracy
  • IoT / IIoT
  • Recommendation Systems
  • Real Time Analytics
  • Google Analytics

If you are learning something by googling these topics, feel free to provide suggestions for adding more words here. You are welcome to discuss / suggest on top of this article as well. Thank you for reading.

Email me: Neil@TechAndTrain.com

Visit my creations:

  • www.TechAndTrain.com
  • www.QandA.in
  • www.TechTower.in

Changes in India’s education system in last few years

  • Institutions of Eminence declared – Complete autonomy given to them
  • University status for IIMs, NITs, IIITs, AIIMS, etc. via Institutions of National Importance route
  • Graded autonomy for UGC affiliated institutions – Based on their accreditation score, they can offer online, distance courses and will have autonomy in academics, faculty recruitment, etc.  
  • Graded autonomy for AICTE affiliated institutions – Based on their accreditation score, they can offer online, distance courses and will have autonomy in academics, faculty recruitment, etc.  
  • MCA shortened to two years from three years – It’s now mapped to a standard university master’s degree of two years 
  • Online degrees approved – Degrees like MBA, MCA, PGDM, etc. are being offered online
  • Rationalization in engineering colleges – Colleges with majority empty seats are being closed with no approvals for new applications by colleges for next few years
  • CGPA system now introduced in almost all universities and colleges
  • Merged single regulator & National Education Policy likely to be finalized in next few months 
  • Executive education programs are getting approvals 
  • Hybrid courses by institutions of eminence & institutes of national importance are starting like Executive MTechs, Executive MBAs, etc. which can be done with your routine job 
  • Foreign collaboration with universities & colleges across the world is becoming easier 
  • Deemed universities with high score in accreditation will not require approvals for open & distance learning courses 

Email me: Neil@TechAndTrain.com

Visit my creations:

  • www.TechAndTrain.com
  • www.QandA.in
  • www.TechTower.in

Three waves of Analytics – Notes on articles by Prof. Davenport

References:

ANALYTICS 1.0 – Business Intelligence, RDBMS & Data Warehousing

  • Vertical scaling
  • Better results and analysis meant higher processing power & memory
  • Complex systems
  • Chances of singular failure
  • Backup was compulsory
  • Storage in RDBMS
  • Transformation in business dimensions and facts in Data Warehouse
  • Descriptive analytics mainly

ANALYTICS 2.0 – BigData, Hadoop, NoSQL & Spark – In memory computing

Problems with Analytics 1.0

  • Costly hardware
  • Large amounts of data
  • Unstructured data

Solution

  • BigData
  • Hadoop – Large files
  • NoSQL – Small files or less size data
  • Horizontal scaling

Problems with BigData

  • Querying unstructured data
  • Large amount of data for real time processing not batch processing

Solution

  • PIG
  • HIVE
  • Spark – In-memory computing
  • Predictive analytics mainly

ANALYTICS 3.0 – Edge Computing, Data Rich Organizations, Real Time Analytics & more

Problems with Analytics 2.0

  • Most analysis was retrospective and for past data
  • Organization wide data also started getting collected but was unused
  • Real time data started to flow in big amounts

Solution

  • Data rich organizations
  • Use data from organization to build products not just mapped to market but also with own organization
  • E.g. Differentiated products in manufacturing to compete with mass economies of scale production
  • Edge computing
  • Real time processing
  • Combined data
  • Embedded analytics
  • Data discovery
  • Cross functional teams
  • Moving to Prescriptive & Real Time analytics

Email me: Neil@TechAndTrain.com

Visit my creations:

  • www.TechAndTrain.com
  • www.QandA.in
  • www.TechTower.in

What should you be doing?

Over the initial years of my experience in Information Technology industry when I worked with various large and medium sized organizations, I programmed, worked in solutions / sales engineering and more. Learned many things across various domains and technologies. Met and built a small network. Then I transitioned to multi-tasking around Education & Information Technology.

Did I face roadblocks, yes – many !!!

  • People
  • Lack of opportunities
  • Lack of resources and more.

Here is what you should be doing if you want to overcome similar roadblocks that you are facing:

  • Build weekly, monthly, yearly and long term goals
  • Have ToDo lists
  • Build priorities & alternatives
  • Plan what you will do and what you will not do
  • Learn to say no and push back on things that don’t resonate with you. Saying yes to everyone does not solve the problem, that will increase your problems
  • Build small steps which you can repeatedly do on daily & weekly basis mapped to small outputs
  • You need to go step by step
  • Many people make a mistake of making a grand goal and struggling on next steps
  • You won’t reach your goals in one shot
  • It’s a step by step journey in which you have to go through with struggle each day & week meeting deadlines, completing work, interacting with people and building your network.
  • Learn to create / write articles, websites, blogs and goals – nothing helps more than building goals and writing them down
  • Learn to give more than you consume especially in knowledge areas. People value others who share and discuss knowledge without expectation. That builds your genuine network which is where your success is. Collaboration and sharing with others to enable their success with no expectations is key to authentic network building
  • Every 3 to 6 months pickup something that you don’t know, learn it, teach it, discuss it, work on it
  • Volunteer for more than what is expected out of you – go an extra mile in your tasks
  • Learn to thank those who helped you on your journey

Email me at: neil@TechAndTrain.com

Visit my URLs:

21st Century Business Management Education: Neither Content nor Pedagogy, Essence is Integration with Triple Bottom Line

Abstract:

In 21st century, many questions clutter our horizon. The way business is passing through sudden and continuous changes; the new business management norms are created every day. The business management education should follow the suit; rather provide lead to the business. Visibly the technological disruptions, social expectations and globalization demands better understanding of its impacts on economy, society and environment in terms of the costs, and benefits. This paper constructs argument emphasizing on core principle of business management education that any course delivery should integrate with triple bottom line. The opportunity costs are enormous if industry or institute fails to do so. Spender rightly poised the questions in his research. What are business schools, and what should they be? What are the social, business, or personal purposes of management education? And how might management education evolve next to meet society’s present needs (J.C. Spender, 2016)? The key questions business school should address revolves around subjects that should be included in syllabus, content of the courses, teaching pedagogy or learning mechanism and recruiting students with right aptitude, attitude and temperament. Business schools attempts to address one or all those ingredients. The more important missing element in business management education is integration of course delivery with the triple bottom line. The existence of business is economic or accounting profit. To be a sustainable business; social and environmental performance of business cannot be ignored. Each course included in the syllabus has a purpose; to help business to enhance triple bottom line. 


Each business problem is unique and one must be able to find solution optimally that fits to the unique situation. There is no one solution which is best to solve a problem. You must be persistent to solve problems on a continuous basis until desired result is obtained. Teacher can guide student, can ignite student’s mind to think beyond horizon. A teacher can expand thinking horizon of the student. In real life situation, a teacher will not accompany student. Student must equip himself to solve the business problem. Therefore, the business management pedagogy seeks involvement of the student while learning. Your solution must be feasible and acceptable to your economic and social surrounding. To solve a problem, you must have information. Scarcity of information is not a problem but abundance information is rather a big challenge. The current age is full of information accessible on public platforms, big challenge is to recognize and extract relevant and reliable information from the information ocean. The next level challenge in this endeavor is to identify real life problems, the application information, and to solve them efficiently and effectively. Learning is a lifelong process. You have to improve your skills on a continuous basis. Without mastering ability to learn new skills, one will become irrelevant. Technology guides you but it also misguides you. It is your ability to judge veracity, relevance, reliability and usefulness of information to churn out the right information.  


Business problems are now seen from prism of economic, social and environmental aspects. Twenty-first-century learning encompass mastery in content producing, synthesizing, and evaluating information from a wide variety of subjects and sources with an understanding of and respect for diverse cultures beside economic understanding. The pillar of success in the 21st century is about knowing how to learn independently. The learning now is eventually be “learner-driven.” The 21st century learning builds upon such past conceptions of learning as “core knowledge in subject areas” and recasts them for today’s world, where a global perspective and collaboration skills are dynamic, critical and focused. It’s no longer enough to “know things”, but to know things to find solution that is economically feasible, socially acceptable and environment friendly. It’s even more important to stay curious about finding out things. We have powerful learning tools at our disposal that allow us to locate, acquire, and even create knowledge much more quickly than our predecessors. Ability to recognize and acquire skills to fit in ever-changing environments is sine-qua-non. No one will tell you what skills are required and the way to acquire it. The self-management is the key to succeed in 21st century. We strongly argue that business management course content should be rich, contemporary, reflecting real life situation, relevant and providing lead to future industry. The course delivery should be learner centered and integrated with triple bottom line.

Vrajlal Sapovadia (Ph.D.)
United States


Ideas on Innovation around Technology. We Thrive On Ideas. We are Learner Centered, Open Source & Digital Focused.