What are
the default configuration files that are used in Hadoop
|
As of
0.20 release, Hadoop supported the following read-only default configurations
-
src/core/core-default.xml
-
src/hdfs/hdfs-default.xml
-
src/mapred/mapred-default.xml
|
How will
you make changes to the default configuration files
|
Hadoop
does not recommends changing the default configuration files, instead it
recommends making all site specific changes in the following files
-
conf/core-site.xml
-
conf/hdfs-site.xml
-
conf/mapred-site.xml
Unless
explicitly turned off, Hadoop by default specifies two resources, loaded
in-order from the classpath:
-
core-default.xml : Read-only defaults for hadoop.
-
core-site.xml: Site-specific configuration for a given hadoop installation.
Hence if
same configuration is defined in file core-default.xml and src/core/core-default.xml then
the values in file core-default.xml (same is true for other 2
file pairs) is used. |
Tuesday, December 4, 2012
Hadoop Interview Questions
Monday, December 3, 2012
Hadoop Interview Question
- What is Hadoop? Brief about the components of Hadoop.
- What are the Hadoop daemon processes tell the components of Hadoop and functionality?
- Tell steps for configuring Hadoop?
- What is architecture of HDFS and flow?
- Can we have more than one configuration setting for Hadoop cluster how can you switch between these configurations?
- What will be your troubleshooting approach in Hadoop?
- What are the exceptions you have come through while working on Hadoop, what was your approach for getting rid of those exceptions or errors?
Thursday, October 4, 2012
100 C Questions and Answers
Tuesday, September 18, 2012
Demystifying Hadoop concepts Series: Safe mode
What is is safe mode of hadoop, may time we come across this exception “ org.apache.hadoop.ipc.RemoteException: org.apache.hadoop.hdfs.server.namenode.SafeModeException” or some other exceptions Which contains safe mode in it
.
First let me tell what Safe mode is in context to Hadoop : as we all know Name node contains fsimage (metadata) of the data present on the cluster, which can be large or small based on the size of the cluster and the size of date present on the cluster, so when the name node starts it loads this fsimage and the edit logs from the disk in the Primary memory RAM for fast processing, and after loading it waits for data nodes to report about the present on those data nodes, so during this process that is loading the fsimage and edit logs and waiting for data nodes to report about the data block in safe mode, which is a read only mode for name node this is done to maintain the consistency of the data present, this is just like saying “ i will not receive any thing till i know what i already have”. And during this period no modification to the file blocks are allowed as to maintain the correctness of the data.
How long safemode exist :
Generally name node automatically comes out of safe mode in 30 seconds if all data present are consistent according the fsimage and the editlogs.
Related commands :
Put Namenode in Safemode: bin/hadoop dfsadmin –safemode
Leave Safemode : bin/hadoop dfsadmin -safemode leave
What to do if you encounter this exception :
First, wait a minute or two and then retry your command. If you just started your cluster, it's possible that it isn't fully initialized yet. If waiting a few minutes didn't help and you still get a "safe mode" error, check your logs to see if any of your data nodes didn't start correctly (either they have Java exceptions in their logs or they have messages stating that they are unable to contact some other node in your cluster). If this is the case you need to resolve the configuration issue (or possibly pick some new nodes) before you can continue.
Tuesday, June 5, 2012
What do you mean by Object Slicing?
{
public int i;
};
class DerivedClass : public BaseClass
{
public int j;
};
int main()
{
BaseClass ObjectOfB;
DerivedClass ObjectOfD;
ObjectOfB = ObjectOfD;
//Here ObjectOfD contains both i and j.
//But only i is copied to ObjectOfB.
}
What is difference between overloading and overriding?
Having same name methods with different parameters is called overloading, while having same name and parameter functions in base and drive class called overriding. |
Tuesday, May 1, 2012
How to : Variable Prameters in c# function.
We can use “params” to enable a method to accept variable number of parameters. For using this we can send comma parameter list as somefunction(1,2,3,1,3,4) to the method.
Sunday, April 22, 2012
Which interface needs to be implemented to create Mapper and Reducer for the Hadoop?
org.apache.hadoop.mapreduce.Mapper ( and ) org.apache.hadoop.mapreduce.Reducer
Thursday, April 19, 2012
how to calculate median in Hive
percentile(BIGINT col, p)
and set p to be 0.5
Will calculate median :)
Tuesday, April 17, 2012
What is the difference between HDFS and NAS ?
- The Hadoop Distributed File System (HDFS) is a distributed file system designed to run on commodity hardware. It has many similarities with existing distributed file systems. However, the differences from other distributed file systems are significant. Following are differences between HDFS and NAS
- In HDFS Data Blocks are distributed across local drives of all machines in a cluster. Whereas in NAS data is stored on dedicated hardware.
- HDFS is designed to work with Map Reduce System, since computation are moved to data. NAS is not suitable for Map Reduce since data is stored separately from the computations.
- HDFS runs on a cluster of machines and provides redundancy using replication protocol. Whereas NAS is provided by a single machine therefore does not provide data redundancy.